MIMIC-MJX: Neuromechanical Emulation of Animal Behavior
Abstract
The primary output of the nervous system is movement and behavior. While recent advances have democratized pose tracking during complex behavior, kinematic trajectories alone provide only indirect access to the underlying control processes. Here we present MIMIC-MJX, a framework for learning biomechanically grounded neural control policies from kinematics. MIMIC-MJX provides a platform for modeling the generative process of motor control by training neural controllers that learn to actuate biomechanical animal models in physics simulation to reproduce real kinematic trajectories. We demonstrate that our implementation is accurate, fast, and generalizable to diverse animal body models, and that it can be trained with modest amounts of motion data. MIMIC-MJX can be used to model motor control policies and simulate behavioral experiments, illustrating its potential as an integrative modeling framework for neuroscience.
Supplementary Video 1: MIMIC-MJX motion capture registration and replay demonstrated across diverse morphologies: rat, fly, mouse arm, worm, and stick insect.
Overview
MIMIC-MJX takes motion capture data and a MuJoCo body model as input, aligns keypoints to the body using stac-mjx, and then trains track-mjx controllers that actuate the model to reproduce the tracked behavior.
MIMIC-MJX is an end-to-end pipeline for training neuromechanical controllers using 3D motion capture data using fast, GPU-accelerated physics simulation.
Starting from pose-tracking data (keypoint trajectories) and a MuJoCo-compatible body model, this framework:
-
Aligns motion capture keypoints to a biomechanically realistic body by performing marker registration and inverse kinematics (IK), resulting in a pose trajectory dataset (stac-mjx)
-
Trains a modular neural network policy using deep reinforcement learning (DRL) to actuate the body so it reproduces the observed behavior in closed-loop simulation (track-mjx).
Together, these stages yield controllers that both match real animal movements and can be repurposed for downstream tasks and neural data analysis.
From kinematics to neuromechanical controllers
Stage 1: Marker registration and inverse kinematics in JAX
stac-mjx is a GPU-parallelized implementation of STAC [1] in JAX. It takes 3D keypoint trajectories and solves for joint angles and derived features of an articulated MuJoCo model, alternating between estimating per-keypoint marker offsets and solving for pose.
The pose solve is a constrained nonlinear least-squares problem: the free root is represented on SE(3), the remaining joint coordinates in Euclidean space, and joint limits enter as inequality constraints. JAX-LS [3] handles those constraints with an augmented-Lagrangian formulation and solves the resulting subproblems using Levenberg–Marquardt with conjugate-gradient linear solves. An optional temporal residual couples jointly solved frames to regularize frame-to-frame configuration changes.
Key features:
- Model-agnostic: use any MuJoCo-compatible body model definition, and configure it with an interactive GUI that maps each tracked keypoint to its corresponding location on the body model.
- High-throughput inverse kinematics: processing rates on the order of hundreds of frames per second on a single GPU, by batching frames through MJX’s vectorized forward kinematics.
The output is a portable HDF5 representation of the full body configuration (joint angles, velocities, and Euclidean coordinates) that becomes the training target for track-mjx.
Stage 2: Deep reinforcement learning for embodied animal motion imitation
track-mjx trains an artificial neural network controller with proximal policy optimization (PPO) [2]. The policy observes the simulated body state and a window of future reference poses, expressed relative to the agent’s current pose, and outputs actions (torques or muscle activations) that drive the model to track the reference trajectory.
The controller is structured as an encoder-decoder:
- The encoder compresses future reference joint angles into a low-dimensional, regularized “motor intention” latent space.
- The decoder combines this intention with current sensory inputs to compute actions.
The latent “motor intention” space, in addition to compressing the representations in a dimensional bottleneck, is regularized using both KL-divergence to a standard Gaussian and AR(1) to improve controllability during decoder reuse in downstream tasks.
MIMIC-MJX enables neuromechanical behavioral analysis and experiment simulation
We explore use cases for MIMIC-MJX in two settings: reuse of the trained controller on a new task as a low-level controller, and analysis of the network activity. These are meant to demonstrate the effectiveness of the trained network for grounding task learning to natural behavior, and the probe-ability of these trained controllers, respectively.
Supplementary Video 2: Transfer learning performance in the rat Bowl Escape task, comparing agents with and without pretrained decoder initialization.
MIMIC-MJX transfer: policy reuse on downstream tasks
The decoder portion of our network, which takes as input a latent vector representing the intended control output, can be reused as a “low-level controller” while a newly initialized high-level policy learns a downstream task by steering it through the motor intention space. We compare this MIMIC-MJX transfer approach against a policy trained from scratch on two rat tasks: Maintain Velocity, where the agent is rewarded for matching a target locomotor speed, and Bowl Escape, where it navigates uneven bowl-shaped terrain and is rewarded for moving away from the center of the bowl as quickly as possible. Both have deliberately naive task designs, with no hand-engineered reward shaping or termination criteria to highlight the advantage that using a low-level controller can offer.
The from-scratch policy learns Maintain Velocity but fails Bowl Escape entirely, while MIMIC-MJX transfer solves both. The behavioral prior also constrains the transfer policy: it tops out at ~0.6 m/s, roughly two standard deviations above the mean locomotor speed in the reference data, whereas the from-scratch policy tracks target speeds all the way to 2.0 m/s, well outside the range represented in that data.
Looking at the solution kinematics makes the difference concrete. On Maintain Velocity, the transfer policy’s footfall timing closely resembles empirically recorded rat gait, while the from-scratch policy settles into an unnatural, jittery pattern. On Bowl Escape, 99.8% of the transfer policy’s frames fall inside the manifold of the recorded reference kinematics, against 2.5% for the from-scratch policy, which occupies an almost entirely separate region of pose space.
PCA visualizations of task trajectories in neural activity space
Here we plot the top 3 principal components of neural activity at different layers in the network for both mouse arm reaching movement and fly walking movement. For the mouse arm, we show how the latent intention vector evolves through the decoder MLP: the intention space is strongly compressed (43.4% of variance on PC1 alone), expands in the early decoder layers, and contracts again by the final layer, with a toroidal structure emerging in the later layers that represents the reaching cycle. For the fly, the intention space forms a cyclic gait manifold: the salient cycles lie along PCs 2 and 3, while gait velocity is represented linearly along the orthogonal PC1 axis. Both the mean position along PC1 and the frequency of the PC2/PC3 cycle are strong predictors of the fly’s locomotor speed (R² = 0.96 and 0.95).
These examples illustrate the possibilities with such neural controllers: analyses built for neural recordings can be run on a fully observable network performing closed-loop control of a biomechanical body.
Getting started with demo notebooks
To get started with MIMIC-MJX, we provide interactive Jupyter notebooks that demonstrate the full pipeline:
- stac-mjx: Try the stac-mjx demo to see how motion capture keypoints are aligned to biomechanical body models.
- track-mjx: Try the demos for the rat, fly, and mouse arm to perform kinematic replay with trained controllers.
These notebooks walk through complete examples using real motion capture data and can be run locally or in cloud environments with GPU support.
Conclusion
MIMIC-MJX lowers the barrier to integrating biomechanical body models, high-resolution behavior, and neural data. By learning controllers that are both biomechanically plausible and experimentally grounded, the framework opens the door to:
- Systematic comparisons between neural representations and learned controllers.
- In silico perturbation experiments targeting specific muscles, joints, or neural populations.
- Extending neuromechanical modeling to new species, tasks, and neural architectures.
BibTeX citation
@misc{mimicmjx2025, title={MIMIC-MJX: Neuromechanical emulation of animal behavior}, author={Zhang, Charles Y. and Yang, Yuanjia and Sirbu, Aidan and Abe, Elliott T. T. and Wärnberg, Emil and Leonardis, Eric J. and Aldarondo, Diego E. and Lee, Adam and Prasad, Aaditya and Foat, Jason and Bian, Kaiwen and Park, Joshua and Bhatt, Rusham and Patel, Vyom N. and Saunders, Hutton and Barbano, Austin and Nagamori, Akira and Thanawalla, Ayesha R. and Huang, Kee Wui and Plum, Fabian and Beck, Hendrik K. and Flavell, Steven W. and Labonte, David and Richards, Blake A. and Brunton, Bingni W. and Azim, Eiman and Ölveczky, Bence P. and Pereira, Talmo D.}, year={2025}, eprint={2511.20532}, archivePrefix={arXiv}, primaryClass={q-bio.NC}}References
[1] Wu, T., Tassa, Y., Kumar, V., Movellan, J. & Todorov, E. STAC: Simultaneous tracking and calibration. 2013 13th IEEE-RAS International Conference on Humanoid Robots (Humanoids) 469–476 (2013).
[2] Schulman, J., Wolski, F., Dhariwal, P., Radford, A. & Klimov, O. Proximal Policy Optimization Algorithms. arXiv:1707.06347 (2017).
[3] Yi, B. jaxls: Nonlinear least squares in JAX (2024).