How We Won IEEE ICRA Robotics Challenges Two Years in a Row 🏆
Over the past two years, I have led my team to win two IEEE ICRA global robotics challenges in two very different fields.
The first was the ICRA 2025 Earth Rover Challenge 2, a global competition in autonomous navigation. The organizers deployed low-cost rovers across seven countries, and we had to operate them remotely while they completed kilometer-scale navigation with no human intervention. The hardest part was not any single algorithm, but the extreme demand for generalization. The robot had very limited sensors—mainly an RGB camera and highly noisy GPS—yet it had to handle different countries, roads, lighting conditions, terrains, and urban structures.
The second was the ICRA 2026 REAL-I Challenge, a manipulation competition much closer to real industrial deployment. The tasks included logistics sorting, loading and unloading automotive parts, and handling daily chemical products. We had to train dual-arm or humanoid manipulation policies from organizer-provided data and then execute the tasks reliably in real environments. Its core constraints were equally clear: we could not collect unlimited new data, we could not validate only in simulation, and the final system had to work consistently on a real robot.
The two competitions could hardly have been more different. One involved long-distance navigation in open environments; the other involved dual-arm manipulation in real industrial settings. One emphasized perception, planning, and robustness; the other emphasized data, policy learning, and generalization. Looking back, however, what enabled us to win two years in a row was not a particular model or an unusually novel method. It was two more fundamental judgments:
- Start from the problem, not from the method.
- Once the problem is clear, use first principles to decompose it into the most basic objectives, constraints, and bottlenecks before choosing a method.
These ideas sound simple, but they matter enormously in real robotics. Robot systems often fail not because they lack the newest model, but because the team never identified the true bottleneck. Even an advanced method creates little value if it does not address the most important constraint in the task. Conversely, a less fashionable combination can outperform a more complex end-to-end system when it precisely matches the problem.
1. Start from the Problem, Not the Method
Robotics now offers an enormous range of methods. Manipulation has WAM, VLA, Diffusion Policy, and many others. Navigation has VLN-style end-to-end models such as NoMaD, SLAM-based mapping and navigation, and optimization-based planning. Every direction contains worthwhile research. But real competitions and real deployment are not method showcases; they are exercises in solving a problem under constraints.
If the first question is “Should I use a VLA?”, “Should I use Diffusion Policy?”, or “Should I use SLAM?”, the order is already backwards. A better starting point is: What does the task actually require the robot to accomplish? What are the hard constraints? Where is failure most likely? Which variables truly determine the final performance?
Method selection becomes meaningful only after the problem is understood. Otherwise, it is easy to choose an advanced method that solves something other than the current bottleneck. A model that excels on a benchmark is not necessarily suited to real industrial manipulation with limited data. A complete SLAM pipeline in a standard setting may not support kilometer-scale navigation on low-cost hardware with monocular RGB, noisy GPS, and open roads.
The first step must therefore be problem-oriented: resist choosing a model too early and decompose the task itself. The method should be the result of problem analysis, not a prior assumption imposed before it.
2. Once the Problem Is Clear, Decompose It with First Principles
Problem orientation answers “Where do we start?” First-principles reasoning answers “How do we break it down?”
In robotics, first principles are not an abstract slogan. They mean reducing the problem to its most basic objective, inputs, constraints, and error sources. What is the system ultimately optimizing? What information is available? Which variables are uncontrollable? Where do the dominant errors come from? What must the model learn? Which components must generalize, and which must remain stable and controllable?
Once the problem reaches this level of decomposition, many method choices become straightforward. You are no longer making a subjective choice among fashionable approaches; you are asking whether a method directly addresses the decisive bottleneck in the task.
3. REAL-I: A Generalizable Manipulation Policy Under Limited Data
Consider the REAL-I Challenge. A method-first approach immediately asks whether to use a VLA, Diffusion Policy, or direct fine-tuning of a large model. A problem-first view reveals something more basic: this is a limited-data learning problem. The organizers provide the data, and without freely collecting more, we must learn a manipulation policy for real industrial scenes that executes reliably on a physical robot.
First-principles decomposition shows that the key bottleneck is not the name of the model. It is three more fundamental issues: data quality, information density, and the balance between model capacity and generalization.
First, data quality. Real-robot data—especially human demonstrations—is naturally noisy. It may contain trajectories that are too long or too short, unstable operations, inconsistent actions, borderline successes, or demonstrations that technically complete the task but follow a process the model should not imitate. Without filtering, the model may memorize noise rather than learn task structure. In a limited-data setting, low-quality samples are especially damaging: they consume training capacity and distort the learned action distribution. Data filtering, trajectory analysis, and quality control are therefore not supporting chores; they are part of the core problem.
Second, information density. The competition requires the use of organizer-provided data, so performance cannot be improved simply by collecting more. But a fixed amount of data does not imply a fixed amount of usable information. If an RGB trajectory is used only for action imitation, it provides one layer of supervision. If it is combined with depth alignment, geometric priors, or structural constraints supplied by other models, the same trajectory can teach richer spatial relationships and task structure. When data quantity is fixed, the real optimization target is the amount of information the model can extract from each sample.
Third, the balance between capacity and generalization. A small model may fail to absorb complex visual-action mappings. An oversized or weakly constrained model may overfit limited data. This is where real robotics differs most sharply from simulation: methods that look similar in simulation can diverge dramatically on a real system because deployment introduces visual shift, object-pose error, execution error, and long-horizon error accumulation. The important question is not which method is most popular, but whether it has enough capacity to learn the task while retaining generalization under limited data.
Our technical decisions in REAL-I therefore revolved around these bottlenecks rather than around using the newest model. Which samples harm training and should be removed? How can we increase the information density of the same data? How can a pretrained model adapt to an industrial task without losing its original generalization? How do we prevent it from memorizing training trajectories and instead make it learn stable spatial and task structure?
This is the relationship between problem orientation and first principles. We first recognize the competition as a real-robot generalization problem under limited data. We then decompose it into data quality, information density, and model generalization. Every method choice follows from those variables.
4. ERC: Generalization in Open Environments
The ICRA 2025 Earth Rover Challenge 2 was a very different case. It looked like a navigation task, but “navigation” alone was not a precise enough description. At a deeper level, it was a reliable-decision problem with weak sensing, long distances, and open environments.
The robot was inexpensive and had very limited sensing—mainly an RGB camera and noisy GPS—yet it had to navigate kilometers through real environments across seven countries. Success depended on more than path planning. The system had to make sufficiently reliable local decisions over time despite incomplete perception, unstable localization, and major environmental variation.
A method-first view suggests two common routes.
The first is end-to-end learning. A visual-navigation model such as NoMaD predicts navigation behavior directly from images. This approach is attractive and consistent with the trend toward learning-based robotics. But from first principles it assumes that we have enough high-quality data, covering a sufficiently broad environmental distribution, to learn a strongly generalizable navigation policy. In this competition, the data was noisy and the environments varied dramatically, so the assumption did not hold. Our experiments confirmed it: directly training or transferring an end-to-end navigation model was unstable under the actual constraints.
The second is classical SLAM + planning. This is a natural navigation pipeline: build a map, localize, then plan. But it also has prerequisites—reliable geometric sensing, usable map quality, and low enough latency. Those conditions were difficult to meet with a low-cost robot and monocular RGB input. After extensive experiments, we found that system latency hurt real-time decisions, while maps built from a monocular camera in open environments were not reliable enough for stable planning. The pipeline was theoretically complete, but its underlying assumptions did not match the task.
We ultimately chose a combination better aligned with the constraints: a learned model for traversability prediction, coupled with a classical optimization-based planner. The learned component did not take responsibility for the entire navigation policy. It answered a more basic question: given the current visual input, which regions are likely traversable, and which are risky? The planner then made stable, controllable, and interpretable decisions over that intermediate representation.
This was neither purely learned nor purely classical, but it matched the problem more precisely. With weak sensors, end-to-end policy learning demanded too much data. A fully SLAM-dependent system demanded better geometry. Traversability prediction turned the problem into an intermediate task well suited to learning, while the planner handled longer-term decisions. This was not compromise for its own sake; it was the system design that naturally followed from the constraints.
Real robot systems do not need to choose between learning and classical methods as rival camps. The real question is which part of the system needs learned generalization, and which part needs the stability and controllability of classical methods.
5. Effective Literature Review Builds a Problem Map
Problem-oriented work cannot rely on intuition alone. One essential skill is using literature review to build a map of the problem, rather than merely collecting methods.
Many people read papers linearly: they see a promising method and want to try it; then they see a new model and think it might also apply. For real robotics, this makes it easy for a single method to steer the whole project. A better approach is to structure the problem space first and then compare methods within it.
For navigation, one might begin with categories such as classical SLAM + planning, end-to-end learning, learned traversability + planner, topological navigation, and GPS-based routing. For each category, analyze its assumptions: which sensors does it require? What data quality does it need? What latency can it tolerate? What are its real-world failure modes? Do those assumptions match the current task?
The same applies to manipulation. Do not begin experiments simply because a VLA, Diffusion Policy, or new policy model is popular. First decompose the problem into data quality, data scale, information density, model capacity, generalization, action representation, visual representation, geometric priors, real-world deployment error, and long-horizon stability. Then ask which dimensions each method addresses and where it may fail.
This changes technical selection from “this method looks strong” to “does this method solve the most important bottleneck in this problem?” The former follows trends; the latter solves problems.
My prior work focused mainly on navigation, and REAL-I was my first serious attempt at manipulation. The key to that transition was whether I could quickly use high-quality literature review and problem decomposition to understand a new field’s structure, taxonomy of methods, and decisive bottlenecks. Once the problem map is clear, method selection and experimental priorities become much clearer as well.
6. Explore Quickly, Then Concentrate on the Critical Point
Once the problem map exists, the next step is rapid experimentation. But the goal is not to run every model available on GitHub; it is to test the assumptions that matter.
In ERC, we tried end-to-end learning and classical SLAM + planning. We did not abandon them because they were insufficiently new or exciting. Experiments showed that their underlying assumptions were not robust under our constraints. End-to-end learning required enough data for generalization. SLAM + planning required reliable geometric mapping and low system latency. Neither condition was sufficiently strong in the competition.
The same applied in REAL-I. Different model families were worth trying, but every experiment had to correspond to an explicit hypothesis. Did it use the data better? Did it reduce overfitting? Did it improve real-world stability? Did action quality persist under distribution shift? When a direction failed, the important question was not how many more hyperparameters to tune, but whether the implementation was incomplete or the method’s assumptions were incompatible with the task.
This judgment is critical. Real projects and competitions do not allow unlimited trial and error. If the idea is sound but the implementation is weak, it deserves further refinement. If the foundational assumption is false, the team should cut its losses early. Exploration can be broad at the beginning, but it must converge later. An effective R&D rhythm is to choose the direction from the problem, identify the critical variables through first principles, test the assumptions experimentally, and then drive the decisive component all the way through.
7. Robot Deployment Is Constraint-Driven Problem Solving, Not Method Stacking
After competing in ICRA robotics challenges for two consecutive years—and winning in navigation and manipulation—the strongest lesson for me is that the scarcest skill in robotics is not mastery of one new model. It is the ability to identify the essence of a problem under complex constraints.
New methods still matter. VLA, Diffusion Policy, SLAM, optimization-based planning, world models, WAM, and real-robot RL all deserve serious study. But methods do not create value automatically in a real task. They create value only when they address the decisive bottleneck.
The conclusion has two levels:
- Start from the problem, not the method. Determine what the task truly requires, what the constraints are, and where failure is most likely.
- Decompose the problem from first principles. Clarify the objective, inputs, data, errors, constraints, and controllable variables before selecting a method.