PhysCaP

PhysCaP

Grounding Code-as-Policy Agent with Physics-Informed Exploration

Chen-Yu Lin1* Jing-Wen Chen1* Hsueh-En Chang1 Hung-An Chen1 Sheng-Hsun Chang1 Chi-Pin Huang2 Fu-En Yang2 Min-Hung Chen2 Yi-Ting Chen3 Yu-Chiang Frank Wang2 Shao-Hua Sun1

1National Taiwan University 2NVIDIA Research 3National Yang Ming Chiao Tung University

*Equal contribution

Accepted at the 8th Robot Learning Workshop @ NeurIPS 2026
Paper (Coming Soon) arXiv Code (Coming Soon) Video Generated Code BibTeX
SCROLL Real robot, sped up

Video

Abstract

We present PhysCaP, a Physics-Informed Code-as-Policy agent system for active perception in robotic manipulation. While vision-language-action policies excel at imitating demonstrations, they rely on passive observation and fail to infer latent physical properties critical for manipulation. PhysCaP augments code-as-policy frameworks with a physics-informed exploration layer that enables explicit information-seeking through interaction. Our method introduces training-free physical property extraction modules that estimate object mass and stiffness from robot proprioception without additional sensors. To balance exploration costs and the efficiency of information obtained, PhysCaP employs a multi-agent design: a Planner that decides when to explore and when to stop, and a Prioritizer that filters implausible interactions and ranks the remainder using a heuristic priority score, enabling efficient, targeted exploration. We evaluate PhysCaP on three real-world tabletop manipulation tasks and a simulated task in LIBERO. The results show that existing passive and naive interactive baselines either fail when physical properties are hidden or over-explore, whereas PhysCaP achieves comparable performance with fewer interactions and reduced execution time. Ablation studies further validate the effectiveness of the proposed physical property extraction modules.

PhysCaP teaser: measuring the mass of cans to find the empty one

PhysCaP augments Code-as-Policy (CaP) agents with physics-informed exploration, enabling inference of latent object properties (e.g., mass) via physical property extraction modules to solve manipulation tasks requiring hidden-state estimation, e.g., removing an empty can.

Method

PhysCaP targets manipulation tasks where success depends on latent physical properties that cannot be seen and no dedicated sensing hardware is available. It actively interacts with objects to acquire the missing information, and once sufficient evidence is collected, it uses the measured properties to synthesize a more informed plan.

Click a module in the figure for more information
Overview of PhysCaP: Planner, Prioritizer and Coding Agent with PhysX modules and system belief

Overview of PhysCaP. Given a task requiring latent physical information, the Planner identifies what information is missing from the visual scene and proposes an initial exploration plan. The Prioritizer then refines this plan by filtering implausible actions and reducing redundant exploration. Finally, a code-generation agent produces executable programs that invoke the PhysX modules (get_mass, get_stiffness) to actively measure the properties. The measurements update the system belief, which the Planner consults to decide whether to explore further or execute the task.

PhysX: Physical Property Extraction Modules

VLMs cannot reliably infer hidden physical properties from visual observations alone, and incorrect estimates propagate to later reasoning and planning. PhysCaP therefore lets the robot interact with objects to measure them. The two PhysX modules are training-free and use only robot proprioception and a standard gripper, with no tactile sensors or other added hardware.

Object mass: get_mass

The robot executes a fixed lift trajectory, raising the end-effector 15 cm. It records joint torques first with an empty grasp and then while holding the object. The torque difference \(\Delta\tau = \tau_{\mathrm{loaded}} - \tau_{\mathrm{empty}}\) isolates the object's gravitational contribution:

\( \hat{m} = \dfrac{J_z \cdot \Delta\tau}{g\,\lVert J_z \rVert^2} \)

\(J_z\): end-effector Jacobian projected onto the vertical direction. The estimate does not depend on where the object was placed. Returns the mass in kilograms (about ±30 g noise); a large negative value means the grasp failed, and NaN means no reading was taken.

Object stiffness: get_stiffness

The gripper closes in fine increments until contact, confirmed by a brief backoff test. It then keeps closing while recording jaw displacement \(d\) and gripper force \(F\) up to a target force, and fits a stiffness coefficient by least squares:

\( k = \dfrac{\Delta F}{\Delta d} \)

Mapped to five levels, from Ultra-Soft (1) to Rigid (5). Each object is measured five times and the final level is a majority vote.

Agent Structure

Measuring takes time, so the agent must balance acting with insufficient information against the cost of unnecessary interactions. PhysCaP splits this decision across separate agents: the Planner decides when to explore and when to stop, the Prioritizer decides what to interact with next, and the Coding Agent turns each decision into executable code.

Find the empty can.
Four cans and a wooden tray on the table
Input: task + camera image
1. Planner Agent

“What do I need to know?”

Current candidates:
blue can black can green can red can
2. Prioritizer Agent

“Wait, is there a better way?”

  1. 2appears unopened, very unlikely to be empty
  2. 1appears unopened, lowest priority
  3. 10has a straw, so it is open: more likely empty
  4. 9also has a straw: highly likely empty
3. Coding Agent
Belief Storage
  • no measurements yet
Empty can found!
The robot placing the red can on the tray
Output: code policy executed

1

Planner Agent

“What do I need to know?”

Given the task and the visual scene, the Planner assesses whether the current observations suffice. If not, it identifies the missing physical information and triggers targeted exploration, proposing candidate interactions.

After each measurement it checks the system belief. It stops exploration once there is sufficient evidence and invokes the Coding Agent, avoiding over-exploration.

2

Prioritizer Agent

“Wait, is there a better way?”

The Prioritizer refines the Planner's candidates. Using visual heuristics, it filters out implausible or redundant interactions and ranks the rest by a priority score, with a brief reason for each score.

For example, it checks an opened can before a sealed one, and a dark avocado before a green one.

3

Coding Agent

The Coding Agent writes Python programs that call the PhysX modules and the low-level control APIs (get_object_pose, goto_pose, open_gripper, close_gripper).

Each measurement updates the system belief. Once the Planner decides the evidence is sufficient, the Coding Agent writes the final program that completes the task.

Why separate agents? We also tried merging prioritization into the Planner prompt (PhysCaP-joint). A single call applied visual cues, such as a straw in a can, only in some trials and otherwise fell back to near-exhaustive exploration. A dedicated Prioritizer applies them more consistently (see Results).

Tasks

We evaluate on three partially observable real-world tabletop tasks with a 6-DoF AgileX PiPER arm and a ZED 2i camera, and a simulated task in LIBERO. Each real-world task uses standardized commercial objects and 3D-printed cubes so setups can be reproduced. Across trials we perturb object positions and colors, vary the words used to describe the target, and randomly sample the hidden physical properties.

Task setups with ground-truth mass and stiffness labels

Ground-truth physical properties of the relevant target objects are annotated in each task setup, with mass (g) and/or stiffness level (1 to 5).

Click to see each task

4× speed, waiting time removed
Separate Planner and Prioritizer.

Identify Empty Can

Real world get_mass

The goal is to identify the single empty soda can among four cans and place it on a wooden tray. The scene contains two sealed cans and two open cans with inserted straws.

A sealed can is unlikely to be empty, so an effective agent should treat the two opened cans as the promising candidates and use get_mass to confirm which one is empty, rather than weighing every can on the table.

Interaction trajectories compared across methods
Identify Empty Can: interaction trajectories for each method

Results

We evaluate over 10 trials per task with three metrics: task success rate (SR \(\uparrow\)), Objects Interacted (OI \(\downarrow\), physical exploration actions before task completion, including repeated interactions with the same object), and Execution Time (Time \(\downarrow\), total robot execution time). OI and Time are reported only for successful episodes. All agents use Gemini 3.1 Pro.

Success rates per method and task

Success rates. The vision-only CaP baseline performs poorly on all three tasks. PhysCaP matches CaP+PhysX on Empty Can and Ripe Avocado, and improves on the longer-horizon Pack Grocery, where CaP+PhysX can lose track of prior measurements.

Exploration efficiency: interactions vs execution time

Exploratory efficiency. Number of physical interactions vs. total robot execution time. Lower-left is more efficient; bubble area shows the standard deviation.

MethodIdentify Empty CanPick Ripe AvocadoPack Grocery Bag
OI ↓Time ↓OI ↓Time ↓OI ↓Time ↓
CaP+PhysX4.00396.84 ± 15 s4.00515.61 ± 9 s9.171019.32 ± 376 s
CaP+PhysX+Planner3.75368.40 ± 62 s4.13563.29 ± 76 s7.80853.14 ± 54 s
PhysCaP-joint2.00222.11 ± 2 s2.71384.39 ± 111 s6.50682.30 ± 139 s
PhysCaP (Ours)2.00256.57 ± 62 s2.00300.47 ± 53 s5.50572.79 ± 61 s

PhysCaP achieves the lowest or tied interaction count on all tasks and the shortest execution time on two of three.

Why separate the Planner and the Prioritizer? PhysCaP-joint puts prioritization into the Planner prompt, so one VLM call decides both when to stop and what to interact with next. It acts on visual cues such as a straw in a can in some trials, but otherwise falls back to near-exhaustive exploration. A separate Prioritizer applies those cues more consistently.

BibTeX

@article{lin2026physcap,
  title   = {PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration},
  author  = {Lin, Chen-Yu and Chen, Jing-Wen and Chang, Hsueh-En and Chen, Hung-An and
             Chang, Sheng-Hsun and Huang, Chi-Pin and Yang, Fu-En and Chen, Min-Hung and
             Chen, Yi-Ting and Wang, Yu-Chiang Frank and Sun, Shao-Hua},
  journal = {arXiv preprint arXiv:2608.21031},
  year    = {2026}
}