Johns Hopkins University · Brains, Bots, and Behavior Lab

MotionForesight

Re-purposing Video Models for Future 3D Scene-Flow Prediction

Johns Hopkins University · Brains, Bots, and Behavior Lab

MotionForesight

Re-purposing Video Models for Future 3D Scene-Flow Prediction

* Equal contribution

From only a short monocular video context, MotionForesight predicts where points on a manipulated object will move next in 3D space—without generating future pixels, requiring language, or making any assumptions on object properties.

Abstract

Predict the motion that matters.

Humans can infer how objects are likely to move from passive observation: a cup may be lifted, a drawer may slide, and a lid may rotate shut. Such predictions expose the physical consequences of interaction needed to act in the real world. We study how to learn this anticipation from ordinary monocular videos of human-object interaction. Given a short observed video context, MotionForesight predicts future 3D trajectories for points on the manipulated object. This casts interaction prediction as object-centered 3D motion forecasting for any object including those that are rigid, articulated, or deformable.

Our key insight is that video prediction models already encode rich priors about how objects move during human interactions. We redirect these priors from pixel prediction toward future 3D scene flow. We start from a dense 3D tracker built on a pretrained video model, generate pseudo-ground-truth tracks from complete clips, and train the forecaster using only the observed frames. We replace future RGB and geometry with learned mask latents and train a lightweight adapter to turn the retrospective tracking representation into a forward predictor, while freezing the large video and tracking components. Using just 40k human videos and no auxiliary inputs such as language, MotionForesight generalizes across diverse out-of-distribution objects, environments, viewpoints, and interactions. It also outperforms substantially larger models that use over a million training videos.

Overview

Learn motion from complete videos, then forecast beyond the observed context.

During training, complete interaction clips provide object masks and pseudo-ground-truth 3D tracks extracted from monocular videos, as supervision. At inference time, the forecasting model receives only a short observed context and predicts the object’s future 3D motion.

Overview animation. Real interaction video, object supervision, and 3D track motion are revealed before the figure steps through observed context, masked future slots, and the predicted trajectory field.
Method

Turn tracking into forecasting.

MotionForesight preserves TrackCraft3R’s point-track interface, but hides every future observation. The model must infer how reference-frame object points continue to move based on the short observed context.

Method animation. The model progresses from observed RGB and pointmaps through time-aligned latent slots, the adapted frozen video prior, residual track latents, and the frozen decoder that produces future metric 3D tracks.
Data curation

Human video, lifted into 3D.

Something-Something-V2 interaction clips are processed offline to recover object masks, metric geometry, camera motion, and dense 3D tracks. Future frames supervise the target—but are not input to the forecasting model.

40K
training videos
7 → 22
observed → future frames
320×576
model resolution
  1. 01
    Find the objectlanguage-guided mask extraction and propagation
  2. 02
    Recover geometrymonocular depth, camera motion, aligned pointmaps
  3. 03
    Track complete clipsdense reference-anchored 3D trajectories
  4. 04
    Hide the futuretrain only from observed RGB and geometry
Real data-curation example. For example, the complete video of a pen being pushed off a table provides a frame-0 object query and 66 reference-anchored metric 3D tracks. Frames after T₁ supervise the target, but remain hidden from the forecasting model.
Qualitative results

One interface, many kinds of motion.

Across everyday objects and articulated scenes, MotionForesight predicts spatially coherent future point motion from only the observed video context.

3D TRACK FORECAST

Shoe

3D TRACK FORECAST

Dishwasher

3D TRACK FORECAST

Mug

3D TRACK FORECAST

Spray bottle

3D TRACK FORECAST

Bowl

3D TRACK FORECAST

Stool

3D TRACK FORECAST

Package

3D TRACK FORECAST

Oven door

3D TRACK FORECAST

Office chair

3D TRACK FORECAST

Door handle

3D TRACK FORECAST

Cup

3D TRACK FORECAST

Backpack

3D TRACK FORECAST

Door handle

3D TRACK FORECAST

Keyboard

3D TRACK FORECAST

Cabinet handle

3D TRACK FORECAST

Door

3D TRACK FORECAST

Small container

3D TRACK FORECAST

Plate

3D TRACK FORECAST

Cup and plate

3D TRACK FORECAST

Cup and plate

3D TRACK FORECAST

Dishwasher

3D TRACK FORECAST

Spray bottle and sink

3D TRACK FORECAST

Mug and sink

3D TRACK FORECAST

Refrigerator door

3D TRACK FORECAST

Shoe

3D TRACK FORECAST

Toy cube

3D TRACK FORECAST

Oil bottle

3D TRACK FORECAST

Door knob

3D TRACK FORECAST

Washer lid

3D TRACK FORECAST

Plush seal

3D TRACK FORECAST

Plush seal

3D TRACK FORECAST

Plush pig

3D TRACK FORECAST

Bottle

3D TRACK FORECAST

Remote control

3D TRACK FORECAST

Package

3D TRACK FORECAST

Umbrella

Citation
@article{bharadhwaj2026motionforesight,
  title={MotionForesight: Re-purposing Video Models for Future 3D Scene-Flow Prediction},
  author={Homanga Bharadhwaj and Yash Jangir},
  year={2026}
}