Skip to content

Embodied AI Field Map

中文版

Embodied AI is a closed-loop systems field, not a single model family. This map separates capability coverage from repository evidence so that a topic name never implies a reproduced result.

Capability stack

Capability Engineering question Best foundation Pipeline contract
Sensing & calibration Are observations synchronized, calibrated, and healthy? Perception and sensors Perception & state estimation
State estimation What state is safe to use, with what uncertainty? Probability and optimization Perception & state estimation
Data & simulation How are tasks, demonstrations, splits, and perturbations produced? MuJoCo · Dataset and training Simulation & data
Policy learning How does observation and language become action? Deep learning · Transformers VLA policy
Predictive models What will happen after an action? Probability and optimization World model planning
Interactive improvement How should behavior improve from reward and failure? Control RL post-training
Generalist policies How are models adapted across datasets and bodies? Dataset and training RFM & cross-embodiment
Task reasoning How is a long instruction decomposed and replanned? Robot systems and safety Embodied reasoning
Manipulation & dexterity How are geometric or learned commands mapped to constrained motion? FK, Jacobian, and IK Dexterous retargeting
Navigation & locomotion How does an embodiment move while remaining localized and stable? Control · Systems and safety Navigation & locomotion
Transfer & deployment What must pass before risk is increased? Evaluation and reproducibility Sim-to-Real

Evidence today

Level Tracks Meaning
Smoke-tested Simulation/data, VLA, world model, RL, dexterous retargeting A lightweight repository path completes; performance is a separate claim.
Interface-tested RFM/cross-embodiment, embodied reasoning Local schemas, adapters, or planners connect without proving real weights or hardware.
Documented Sim-to-Real, perception/state estimation, navigation/locomotion The engineering contract and gates exist; no universal local command represents the system.

Choose by research goal

Goal Start Then prove
Learn robot learning from zero Foundations overview Complete one smoke-tested pipeline and retain its artifacts.
Build a multimodal policy VLA pipeline Closed-loop success, language ablation, latency, and failure cases.
Study prediction and planning World-model pipeline Multi-step rollout error and planned task success separately.
Work across robot bodies RFM pipeline Action semantics, adapter coverage, and per-embodiment results.
Study dexterous hands Retargeting pipeline Geometry, temporal quality, contact/task evidence, and hardware evidence separately.
Build mobile or legged systems Navigation/locomotion contract Localization, tracking, collision/fall, recovery, and transfer evidence.

Deliberate non-claims

The repository does not currently claim a reproduced SLAM benchmark, navigation success rate, legged-locomotion policy, general-purpose hardware deployment, or competitive large-scale foundation-model result. These are visible expansion targets, not hidden assumptions.