Research
My work sits at the intersection of optimization, control and machine learning for robotics. I am interested in methods that let a robot exploit the structure of a problem and prior knowledge, learn from limited data, produce safe and efficient behaviors, and adapt when the model, the environment, or the robot itself changes.
Evolutionary algorithms & Quality-Diversity
Evolutionary algorithms are useful when the search space is large, non-linear, or offers no reliable gradients. In robotics, however, a single solution with the best objective value is often not enough. A robot that explores, moves through an unknown environment, or has been damaged needs alternative behaviors, so that it can pick one that still works in its current situation. This is the motivation behind Quality-Diversity (QD) methods: they build a repertoire of solutions that combine high performance with meaningful behavioral diversity.
My work on CVT-MAP-Elites looked at scaling MAP-Elites-style methods to high-dimensional behavior spaces. More recently, VQ-Elites learns the structure of the behavior space directly from data, removing the need to hand-design behavior descriptors. The same idea of hierarchical, reusable structure underlies Hierarchical Trial & Error, aimed at damage recovery on physical robots, and the EvoDSM framework for optimizing doubly-stochastic matrices with swarm and evolutionary algorithms. Today, QD is also central to the autonomous, open-ended skill discovery work in the NOSALRO project, where it is combined with simulators, variational autoencoders, and reinforcement learning.
Optimization, optimal control & numerical methods
In optimal control and trajectory generation, a method's value depends on whether it can solve the problem within the available time budget. I care about exploiting the algebraic and computational structure of these problems — sparsity of the linear systems, derivatives, constraints, and the geometry of the state — so that trajectory optimization and Model Predictive Control (MPC) can run in real time.
Along these lines, we are developing SSQP, a structure-exploiting Sequential Quadratic Programming solver for nonlinear optimal control problems, with native support for manifold-valued states such as floating-base orientations. In AHMP, contact-sequence discovery via a Mixed-Distribution Cross-Entropy Method is combined with whole-body trajectory optimization in the tangent space of SE(3). Our comparative study of floating-base parameterizations looks at how the choice between Euler angles, quaternions, and tangent-space representations affects convergence and motion quality, and the 3D-orientations benchmark extends this question to learning and optimization more broadly. On the constrained-optimization side, UPSO-QP and its successor, the Hybrid Augmented Lagrangian (HyAL) method, embed evolutionary search inside an Augmented Lagrangian loop to combine gradient-free robustness with fast, precise convergence.
Data-efficient robot learning
Learning on real robots is constrained by time, hardware wear, safety, and the cost of failed trials. I am interested in methods that learn robot skills efficiently — combining simulators, dynamics models, prior knowledge, and suitable policy representations. This direction is summarized in my survey on policy search algorithms for learning robot controllers in a handful of trials.
Black-DROPS uses black-box optimizers for model-based policy search, while parameterized black-box priors scale this up to high-dimensional robots without requiring the prior to be exact. Choosing among several available priors is studied in MLEI, and robust reinforcement learning with ALOQ addresses robust learning when simulation and reality differ.
Adaptation, robustness & safe behavior
A robot operating outside the lab has to cope with a changing environment, unfamiliar situations, and sensor or actuator damage. Reset-free Trial-and-Error studies how a robot's locomotion abilities can be restored after mechanical damage, without human intervention and without resetting the robot after every episode — a direction closely tied to the use of Quality-Diversity repertoires, which offer alternative behaviors when the original solution is no longer feasible.
Safety, in turn, should not be an afterthought. I am interested in combining constraints, probabilistic models, and controllers that estimate uncertainty and adjust their behavior accordingly. Self-correcting QP control is one example of this, while our bimanual manipulation and human-to-robot handover benchmarks support reproducible evaluation of such methods.
Machine learning, RL & behavior representations
I am interested in representations that make learning more efficient, interpretable, and stable. Behavior Policy Learning combines solution sketches, model-based controllers, and simulation to learn multi-stage tasks, while our work on Behavior Trees for movement skills looks at a structured, reusable policy representation.
A second thread combines neural models with dynamical systems and controllers. Autonomous Neural Dynamic Policies pair the expressiveness of neural networks with stability guarantees, and in PGTT terrain perception and gait structure are built into the training of a perceptive legged-locomotion RL controller, without constraining the action space through oscillators or inverse kinematics. The IBRICS project studies vision and machine learning for imitating human behavior in robotic systems, and this direction also covers imitation learning, diffusion models, normalizing flows, and Bayesian inference for perception, planning, and control.
Simulation, software & reproducible robotics
Simulation and open-source software are essential for fast iteration and reproducible research. RobotDART, built on the DART physics engine, provides an efficient simulator for robotics and machine-learning researchers, and Limbo supports Gaussian Processes and data-efficient optimization. SSQP extends this philosophy to optimal control, aiming for tools usable both in research experiments and in real-time applications.
Overall, I am interested in bridging evolutionary search, mathematical optimization, and machine learning: exploiting structure and prior knowledge to reduce the amount of data required, using safe and robust controllers to transfer solutions from simulation to the real world, and building robots that do not simply execute one fixed policy, but can explore, learn, and adapt.
For the complete, up-to-date list of papers, book chapters and preprints, see the Publications page.