The core idea
Multimodal large language models can integrate long visual histories, reason under partial observability, and learn from examples. Vision-language-action models, however, typically inherit pretrained representations without using that contextual capacity as episode memory. Memory-dependent policies instead add task-specific samplers, stores, retrievers, or compressors.
PonderPounce takes a different route. Ponder, a System 2 MLLM, accumulates observations, demonstrations, and prior cognition in its native causal context. Pounce, a System 1 action model, receives the current observation and the newest continuous cognition token together with its age. The two pretrained systems run on decoupled clocks and are trained jointly end to end, without a purpose-built memory module or separate bridge pretraining.
Retains the episode in a pretrained MLLM context and produces fresh continuous cognition at sparse queries.
Routes only the latest cognition token and its age, keeping long-context processing off the action path.
Conditions a fast VLA controller on current sensory input and cognition while playing action chunks back at 20 Hz in simulation.
Architecture
Ponder accumulates the instruction, demonstrations, and observations. Transition tokens gate optional internal demonstration reasoning and subgoal text, while carrier states form continuous cognition. Pounce combines the newest cognition, its age, and the current observation to predict an action chunk. On an H100, optimized per-call p50 latencies are 78 ms for cognition-only refresh and 25 ms for Pounce invocation.
Results
Average success on memory-dependent control and demonstration-conditioned manipulation.
RoboMME average success rate across base-scale, 9×-data, and oracle or human reference settings.
| Method | Episode memory | 1× data | 9× data |
|---|---|---|---|
| π0.5 | None | 17.93% | — |
| π0.5 + past actions | Action history | 19.73% | — |
| SAM2Act+ | SAM2 bank + attention | 21.37% | — |
| SimpleSG+QwenVL | Subgoal-text list | 29.00% | — |
| GroundSG+QwenVL | Subgoal-text + box lists | 32.70% | — |
| MemER | Selected-keyframe image buffer | 42.38% | — |
| FrameSamp+Modul | Sampled-frame token buffer | 44.51% | 57.88% |
| PonderPounce (0.8B Ponder) | Native MLLM context | 54.12% | — |
| PonderPounce (9B Ponder) | Native MLLM context | 60.83% | 75.54% |
| SimpleSG+Oracle | Oracle subgoal | 49.58% | |
| GroundSG+Oracle | Grounded oracle subgoal | 84.08% | |
| Human | Human memory · reference | 90.50% | |
Real-world evaluation
Three methods trained on the same 710 teleoperated episodes were tested on 59 held-out scenes. PonderPounce reaches 60.98% mean success across four tasks, versus 40.67% for FrameSamp+Modul and 23.99% for π0.5.
| Method | PutFruits | TrackCube | RepickBlock | DrawPattern | Mean |
|---|---|---|---|---|---|
| π0.5 | 9/14 | 4/16 | 0/14 | 1/15 | 23.99% |
| FrameSamp+Modul | 7/14 | 5/16 | 3/14 | 9/15 | 40.67% |
| PonderPounce | 11/14 | 10/16 | 6/14 | 9/15 | 60.98% |
Entries show successful trials / total trials. The mean is the unweighted average of the four task success rates.
Real robots in action
Three examples per task, with three policies compared on each evaluation scene.
Selected qualitative examples; success and failure labels are from the evaluation records. Aggregate results are reported above.
Red border: task demonstration, followed by the policy rollout.
Memory in action
In RoboMME PickHighlight, target markers appear briefly and then disappear. A current-observation policy searches and times out; PonderPounce retains the earlier cue in episode context and grasps both highlighted targets.
Cognition analysis
Delivering fresh cognition only at predicted transitions drops RoboMME success from 60.83% to 1.83% when held cognition is presented as 300 ms old. Reporting its true age improves success to 22.42%, still below regular refresh. This shows the importance of both timely cognition and an accurate age signal.
Offline staleness tests show the same tradeoff: frequent-refresh checkpoints fit fresh cognition best, while models trained with slower refresh intervals are more robust when the cognition grows old.
Normalized validation loss under stale cognition. Lower is better; bands show 95% task-stratified bootstrap confidence intervals.
Citation
@article{choi2026ponderpounce,
title={PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control},
author={Choi, Suhwan and Jung, Jaeyoon and Kim, Sungkyung and Lee, Yunsung and Yu, Youngjae},
journal={arXiv preprint arXiv:2608.24115},
year={2026}
}