PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control

1 MAUM.AI 2 Seoul National University 3 Georgia Institute of Technology
* Equal contribution   † Co-corresponding authors
Under Review
PonderPounce overview with a real-robot TrackCube example, system latency, and RoboMME results

Ponder remembers and reasons over the episode. Pounce acts quickly from the latest cognition.

The core idea

Use pretrained context as robot memory

Multimodal large language models can integrate long visual histories, reason under partial observability, and learn from examples. Vision-language-action models, however, typically inherit pretrained representations without using that contextual capacity as episode memory. Memory-dependent policies instead add task-specific samplers, stores, retrievers, or compressors.

PonderPounce takes a different route. Ponder, a System 2 MLLM, accumulates observations, demonstrations, and prior cognition in its native causal context. Pounce, a System 1 action model, receives the current observation and the newest continuous cognition token together with its age. The two pretrained systems run on decoupled clocks and are trained jointly end to end, without a purpose-built memory module or separate bridge pretraining.

Ponder · System 2

Retains the episode in a pretrained MLLM context and produces fresh continuous cognition at sparse queries.

Asynchronous interface

Routes only the latest cognition token and its age, keeping long-context processing off the action path.

Pounce · System 1

Conditions a fast VLA controller on current sensory input and cognition while playing action chunks back at 20 Hz in simulation.

Architecture

Context accumulates; control stays fast

PonderPounce architecture showing instructions, demonstrations, and observations entering Ponder, with cognition tokens routed to Pounce

Ponder accumulates the instruction, demonstrations, and observations. Transition tokens gate optional internal demonstration reasoning and subgoal text, while carrier states form continuous cognition. Pounce combines the newest cognition, its age, and the current observation to predict an action chunk. On an H100, optimized per-call p50 latencies are 78 ms for cognition-only refresh and 25 ms for Pounce invocation.

Results

Pretrained context transfers to control

Average success on memory-dependent control and demonstration-conditioned manipulation.

RoboMME average success rate comparison across baselines, PonderPounce, oracle methods, and humans.

RoboMME average success rate across base-scale, 9×-data, and oracle or human reference settings.

60.83%
RoboMME success
with base-scale data
75.54%
RoboMME success
with 9× data
6.71pp
Gain from scaling Ponder
from 0.8B to 9B
60.98%
Real-robot mean success
across four tasks
Method Episode memory 1× data 9× data
π0.5 None 17.93% —
π0.5 + past actions Action history 19.73% —
SAM2Act+ SAM2 bank + attention 21.37% —
SimpleSG+QwenVL Subgoal-text list 29.00% —
GroundSG+QwenVL Subgoal-text + box lists 32.70% —
MemER Selected-keyframe image buffer 42.38% —
FrameSamp+Modul Sampled-frame token buffer 44.51% 57.88%
PonderPounce (0.8B Ponder) Native MLLM context 54.12% —
PonderPounce (9B Ponder) Native MLLM context 60.83% 75.54%
SimpleSG+Oracle Oracle subgoal 49.58%
GroundSG+Oracle Grounded oracle subgoal 84.08%
Human Human memory · reference 90.50%
Same controller, larger context engine: replacing pretrained Qwen3.5-0.8B with 9B raises RoboMME success from 54.12% to 60.83%, a 6.71 percentage-point gain without changing the Pounce architecture or interface. In a separately trained, matched-supervision control, removing execution history lowers 9B PonderPounce success from 60.83% to 26.21%.

Real-world evaluation

Real-robot results

Three methods trained on the same 710 teleoperated episodes were tested on 59 held-out scenes. PonderPounce reaches 60.98% mean success across four tasks, versus 40.67% for FrameSamp+Modul and 23.99% for π0.5.

Method PutFruits TrackCube RepickBlock DrawPattern Mean
π0.5 9/14 4/16 0/14 1/15 23.99%
FrameSamp+Modul 7/14 5/16 3/14 9/15 40.67%
PonderPounce 11/14 10/16 6/14 9/15 60.98%

Entries show successful trials / total trials. The mean is the unweighted average of the four task success rates.

Trossen ALOHA Stationary AI Kit used for the four real-robot tasks
Trossen ALOHA Stationary AI Kit. Only the right arm is used.

Real robots in action

Real-robot task examples

Three examples per task, with three policies compared on each evaluation scene.

Example:

Selected qualitative examples; success and failure labels are from the evaluation records. Aggregate results are reported above.

Red border: task demonstration, followed by the policy rollout.

Memory in action

Acting after the evidence disappears

RoboMME PickHighlight example where PonderPounce remembers transient target markers and succeeds while the baseline times out

In RoboMME PickHighlight, target markers appear briefly and then disappear. A current-observation policy searches and times out; PonderPounce retains the earlier cue in episode context and grasps both highlighted targets.

Cognition analysis

Fresh cognition matters

Delivering fresh cognition only at predicted transitions drops RoboMME success from 60.83% to 1.83% when held cognition is presented as 300 ms old. Reporting its true age improves success to 22.42%, still below regular refresh. This shows the importance of both timely cognition and an accurate age signal.

Offline staleness tests show the same tradeoff: frequent-refresh checkpoints fit fresh cognition best, while models trained with slower refresh intervals are more robust when the cognition grows old.

Normalized validation loss as cognition becomes stale for models trained with one, two, and four second refresh intervals

Normalized validation loss under stale cognition. Lower is better; bands show 95% task-stratified bootstrap confidence intervals.

Citation

BibTeX

@article{choi2026ponderpounce,
  title={PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control},
  author={Choi, Suhwan and Jung, Jaeyoon and Kim, Sungkyung and Lee, Yunsung and Yu, Youngjae},
  journal={arXiv preprint arXiv:2608.24115},
  year={2026}
}