MIMIC-MJX

Preprint
Affiliations
1 Department of Organismic and Evolutionary Biology, Harvard University, Cambridge, MA, USA
2 Computational Neurobiology Laboratory, Salk Institute for Biological Studies, La Jolla, CA, USA
3 Neurosciences Graduate Program, University of California San Diego, La Jolla, CA, USA
4 Mila, Montréal, QC, Canada
5 School of Computer Science, McGill University, Montréal, QC, Canada
6 Biology Department, University of Washington, Seattle, WA, USA
7 eScience Institute, University of Washington, Seattle, WA, USA
8 Computational Neuroscience Center, University of Washington, Seattle, WA, USA
9 Department of Brain and Cognitive Sciences, Massachusetts Institute of Technology, Cambridge, MA, USA
10 Picower Institute for Learning and Memory, Massachusetts Institute of Technology, Cambridge, MA, USA
11 Molecular Neurobiology Laboratory, Salk Institute for Biological Studies, La Jolla, CA, USA
12 Department of Bioengineering, Imperial College London, London, United Kingdom
13 Howard Hughes Medical Institute, Cambridge, MA, USA
14 Department of Neurology and Neurosurgery, McGill University, Montréal, QC, Canada
15 Learning in Machines and Brains Program, Canadian Institute for Advanced Research, Toronto, ON, Canada
16 Montréal Neurological Institute, McGill University, Montréal, QC, Canada
17 Center for Brain Science, Harvard University, Cambridge, MA, USA
18 Kempner Institute, Harvard University, Cambridge, MA, USA
19 Neuromatch, Beaverton, OR, USA
*These authors contributed equally., ‡Co-senior authors., §Co-corresponding authors. Emails: [email protected]; [email protected]

MIMIC-MJX: Neuromechanical Emulation of Animal Behavior

Abstract

The primary output of the nervous system is movement and behavior. While recent advances have democratized pose tracking during complex behavior, kinematic trajectories alone provide only indirect access to the underlying control processes. Here we present MIMIC-MJX, a framework for learning biomechanically grounded neural control policies from kinematics. MIMIC-MJX provides a platform for modeling the generative process of motor control by training neural controllers that learn to actuate biomechanical animal models in physics simulation to reproduce real kinematic trajectories. We demonstrate that our implementation is accurate, fast, and generalizable to diverse animal body models, and that it can be trained with modest amounts of motion data. MIMIC-MJX can be used to model motor control policies and simulate behavioral experiments, illustrating its potential as an integrative modeling framework for neuroscience.

Supplementary Video 1: MIMIC-MJX motion capture registration and replay demonstrated across diverse morphologies: rat, fly, mouse arm, worm, and stick insect.

Overview

Diagram of the MIMIC-MJX pipeline from motion capture to neuromechanical policy.

MIMIC-MJX takes motion capture data and a MuJoCo body model as input, aligns keypoints to the body using stac-mjx, and then trains track-mjx controllers that actuate the model to reproduce the tracked behavior.

MIMIC-MJX is an end-to-end pipeline for training neuromechanical controllers using 3D motion capture data using fast, GPU-accelerated physics simulation.

Starting from pose-tracking data (keypoint trajectories) and a MuJoCo-compatible body model, this framework:

  1. Aligns motion capture keypoints to a biomechanically realistic body by performing marker registration and inverse kinematics (IK), resulting in a pose trajectory dataset (stac-mjx)

  2. Trains a modular neural network policy using deep reinforcement learning (DRL) to actuate the body so it reproduces the observed behavior in closed-loop simulation (track-mjx).

Together, these stages yield controllers that both match real animal movements and can be repurposed for downstream tasks and neural data analysis.

From kinematics to neuromechanical controllers

Stage 1: Marker registration and inverse kinematics in JAX

stac-mjx diagram showing marker registration and inverse kinematics

stac-mjx is a GPU-parallelized implementation of STAC [1] in JAX. It takes 3D keypoint trajectories and solves for joint angles and derived features of an articulated MuJoCo model, alternating between estimating per-keypoint marker offsets and solving for pose.

The pose solve is a constrained nonlinear least-squares problem: the free root is represented on SE(3), the remaining joint coordinates in Euclidean space, and joint limits enter as inequality constraints. JAX-LS [3] handles those constraints with an augmented-Lagrangian formulation and solves the resulting subproblems using Levenberg–Marquardt with conjugate-gradient linear solves. An optional temporal residual couples jointly solved frames to regularize frame-to-frame configuration changes.

Key features:

The output is a portable HDF5 representation of the full body configuration (joint angles, velocities, and Euclidean coordinates) that becomes the training target for track-mjx.

Stage 2: Deep reinforcement learning for embodied animal motion imitation

track-mjx diagram showing deep reinforcement learning for naturalistic sensorimotor control

track-mjx trains an artificial neural network controller with proximal policy optimization (PPO) [2]. The policy observes the simulated body state and a window of future reference poses, expressed relative to the agent’s current pose, and outputs actions (torques or muscle activations) that drive the model to track the reference trajectory.

The controller is structured as an encoder-decoder:

The latent “motor intention” space, in addition to compressing the representations in a dimensional bottleneck, is regularized using both KL-divergence to a standard Gaussian and AR(1) to improve controllability during decoder reuse in downstream tasks.

MIMIC-MJX enables neuromechanical behavioral analysis and experiment simulation

We explore use cases for MIMIC-MJX in two settings: reuse of the trained controller on a new task as a low-level controller, and analysis of the network activity. These are meant to demonstrate the effectiveness of the trained network for grounding task learning to natural behavior, and the probe-ability of these trained controllers, respectively.

Supplementary Video 2: Transfer learning performance in the rat Bowl Escape task, comparing agents with and without pretrained decoder initialization.

MIMIC-MJX transfer: policy reuse on downstream tasks

MIMIC-MJX transfer architecture compared against learning from scratch, with task reward curves and locomotor speed tracking

The decoder portion of our network, which takes as input a latent vector representing the intended control output, can be reused as a “low-level controller” while a newly initialized high-level policy learns a downstream task by steering it through the motor intention space. We compare this MIMIC-MJX transfer approach against a policy trained from scratch on two rat tasks: Maintain Velocity, where the agent is rewarded for matching a target locomotor speed, and Bowl Escape, where it navigates uneven bowl-shaped terrain and is rewarded for moving away from the center of the bowl as quickly as possible. Both have deliberately naive task designs, with no hand-engineered reward shaping or termination criteria to highlight the advantage that using a low-level controller can offer.

The from-scratch policy learns Maintain Velocity but fails Bowl Escape entirely, while MIMIC-MJX transfer solves both. The behavioral prior also constrains the transfer policy: it tops out at ~0.6 m/s, roughly two standard deviations above the mean locomotor speed in the reference data, whereas the from-scratch policy tracks target speeds all the way to 2.0 m/s, well outside the range represented in that data.

Footfall timing rasters and bowl escape kinematic occupancy comparing MIMIC-MJX transfer, from-scratch training, and recorded rat data

Looking at the solution kinematics makes the difference concrete. On Maintain Velocity, the transfer policy’s footfall timing closely resembles empirically recorded rat gait, while the from-scratch policy settles into an unnatural, jittery pattern. On Bowl Escape, 99.8% of the transfer policy’s frames fall inside the manifold of the recorded reference kinematics, against 2.5% for the from-scratch policy, which occupies an almost entirely separate region of pose space.

PCA visualizations of task trajectories in neural activity space

PCA analysis of neural representations in trained policies

Here we plot the top 3 principal components of neural activity at different layers in the network for both mouse arm reaching movement and fly walking movement. For the mouse arm, we show how the latent intention vector evolves through the decoder MLP: the intention space is strongly compressed (43.4% of variance on PC1 alone), expands in the early decoder layers, and contracts again by the final layer, with a toroidal structure emerging in the later layers that represents the reaching cycle. For the fly, the intention space forms a cyclic gait manifold: the salient cycles lie along PCs 2 and 3, while gait velocity is represented linearly along the orthogonal PC1 axis. Both the mean position along PC1 and the frequency of the PC2/PC3 cycle are strong predictors of the fly’s locomotor speed (R² = 0.96 and 0.95).

These examples illustrate the possibilities with such neural controllers: analyses built for neural recordings can be run on a fully observable network performing closed-loop control of a biomechanical body.

Getting started with demo notebooks

To get started with MIMIC-MJX, we provide interactive Jupyter notebooks that demonstrate the full pipeline:

These notebooks walk through complete examples using real motion capture data and can be run locally or in cloud environments with GPU support.


Conclusion

MIMIC-MJX lowers the barrier to integrating biomechanical body models, high-resolution behavior, and neural data. By learning controllers that are both biomechanically plausible and experimentally grounded, the framework opens the door to:

BibTeX citation

@misc{mimicmjx2025,
title={MIMIC-MJX: Neuromechanical emulation of animal behavior},
author={Zhang, Charles Y. and Yang, Yuanjia and Sirbu, Aidan and Abe, Elliott T. T. and Wärnberg, Emil and Leonardis, Eric J. and Aldarondo, Diego E. and Lee, Adam and Prasad, Aaditya and Foat, Jason and Bian, Kaiwen and Park, Joshua and Bhatt, Rusham and Patel, Vyom N. and Saunders, Hutton and Barbano, Austin and Nagamori, Akira and Thanawalla, Ayesha R. and Huang, Kee Wui and Plum, Fabian and Beck, Hendrik K. and Flavell, Steven W. and Labonte, David and Richards, Blake A. and Brunton, Bingni W. and Azim, Eiman and Ölveczky, Bence P. and Pereira, Talmo D.},
year={2025},
eprint={2511.20532},
archivePrefix={arXiv},
primaryClass={q-bio.NC}
}

References

[1] Wu, T., Tassa, Y., Kumar, V., Movellan, J. & Todorov, E. STAC: Simultaneous tracking and calibration. 2013 13th IEEE-RAS International Conference on Humanoid Robots (Humanoids) 469–476 (2013).

[2] Schulman, J., Wolski, F., Dhariwal, P., Radford, A. & Klimov, O. Proximal Policy Optimization Algorithms. arXiv:1707.06347 (2017).

[3] Yi, B. jaxls: Nonlinear least squares in JAX (2024).