Spatial Concept Bottleneck for Interpretable Robot Learning

Anonymous Authors
Under review

Abstract

Robot learning models have been proliferating with increasingly stronger results for both generalist and specialist policies. However, true deployment potential of these learned policies is still unknown. As the community closes in on translational research, it is now ever more crucial to develop mechanisms that not only enable interpretability but intervention as well, e.g., through concept bottleneck models. In this work, we present an initial investigation along the lines of an interpretable spatial concept bottleneck that acts as a meaningful intermediate representation between a perception-planning backbone and an action head. Specifically, we define a spatial concept as a geometric relationship between entities corresponding to the task and the robot, which can be computed directly from the scene's geometry, and instantiate it for both navigation and manipulation. We present a series of experiments demonstrating that such a spatial concept bottleneck maintains original performance (without the bottleneck), while serving as an interpretable interface through which humans or VLMs can intervene to correct or steer the robot.

Navigation: the VLM moves the costmap's minimum off the wall, and the robot reaches the goal.

Manipulation: the source concept moves to the other bowl, and the robot moves that bowl.

Method

We define a spatial concept as a geometric relationship between entities corresponding to the task and the robot, which can be computed directly from the scene's geometry. The perception backend \(g\) predicts the spatial concept \(\hat{c}\) and the actor/controller head maps the concept and optionally the proprioceptive state \(\rho_t\) to action \(\hat{y}_t = f(\hat{c}_t, \rho_t)\). The robot executes \(\hat{y}_t\), observes \(o_{t+1}\), and the loop repeats.

\[ c^k_t = \psi\big(a^k_t, r\big), \qquad r \in R_t \]

The task mandates an exocentric entity \(a^k_t\), such as a navigation sub-goal or a destination point for the gripper, and the robot provides an egocentric entity \(r\).

Navigation

Navigation setup
Spatial concept bottleneck for navigation. The concept predictor \(g\) (VGGT with fine-tuned global attention and a costmap head) predicts the costmap \(\hat C\) from the current view \(o_t\), the eight localized traversal frames \(m_{1..8}\) and the sub-goal \(a_t\). The task predictor \(f\) (the action head) takes only the costmap as input and outputs five waypoints.
\[ C_t(u) = \psi\big(a^1_t, X_t(u)\big) = \big\lVert a^1_t - X_t(u) \big\rVert_2 \]

We define a spatial concept bottleneck for navigation as a visual representation of navigation affordances in the robot's current view. The sub-goal pixel is the only anchor, and the egocentric entity is the 3D point \(X_t(u)\) of each pixel \(u\) of the current observation.

Manipulation

Manipulation setup
Spatial concept bottleneck on SmolVLA. Left: the concept \(c\) is where the source (\(c^1\)) and the destination (\(c^2\)) are relative to the gripper. The frozen vision encoder feeds only the concept predictor \(g\), whose prediction \(\hat c\) enters the task predictor \(f\) as concept tokens, after the instruction \(H\) and the state, and the image tokens are removed. An intervention replaces \(\hat c\) with \(\tilde c\).
\[ c^k_t = \psi (a^k_t, r_t) = \beta\left(a^k_t - r_t\right), \qquad k \in \{1, 2\} \]

For manipulation, we define two exocentric spatial entities: source and destination, corresponding respectively to the pick and place objects. The source object defines where to act as \(a^1_t\) and the destination defines where to place as \(a^2_t\). We use a single egocentric spatial entity that represents the end effector position \(r_t\), and \(\beta\) discretizes each axis of the difference in positions into 13 bins on a logarithmic scale. The concept therefore consists of six values, the \(x\), \(y\) and \(z\) offsets of the source and of the destination from the gripper. The bins are narrow near the gripper and wide far from it, giving centimetre precision close to the object and a coarse direction elsewhere.

The concept along one rollout

The concept along one rollout (LIBERO-Spatial, true concepts). Each box is the bin \(\beta\) puts the source (orange) or the destination (blue) in, drawn around the gripper. Green represents the object's true position. Below each frame, the 13 bins of one axis to scale, with the source's and the destination's bin. The triangle is the gripper.

Results

Navigation

Spatial CBM is on par or better than predicting Direct Actions.

ImitateAlt GoalShortcutReverseAverage
MethodSPLSSPLSPLSSPLSPLSSPLSPLSSPLSPLSSPL
GT-Geodesics96.43±2.3498.78±1.7286.53±4.0791.16±3.2088.42±0.2091.67±1.1982.99±0.3291.97±0.3988.59±1.6393.40±0.83
GT-ED72.16±2.0074.58±1.5370.29±4.8978.49±5.9467.76±2.2870.72±3.1968.41±2.5177.53±1.1269.65±1.7875.33±1.35
Direct Action82.61±5.7786.92±3.0759.99±13.0767.84±12.0761.15±5.9171.86±4.2118.41±3.2834.47±2.8255.54±0.4765.28±0.49
Spatial CAM87.25±0.0389.70±0.4628.18±7.1641.45±9.0138.13±0.2742.01±0.707.86±2.2515.71±0.7840.36±1.1747.22±2.00
Spatial CBM*81.14±5.8386.44±2.5457.81±2.9468.08±3.5666.06±3.9274.20±2.9712.55±0.0827.33±2.2754.39±3.1964.01±1.70
Spatial CBM80.80±2.0886.12±1.8268.27±6.4373.89±5.3373.06±2.6980.29±1.8019.38±10.2233.73±9.1260.38±5.3668.51±4.52

SPL and SSPL on four navigation tasks, with the best non-oracle scores in bold.


Manipulation

With true concepts the spatial bottleneck performs on par with Direct Action, and swapping the source and destination brings the GT policy down to 4.5%.

MethodT1T2T3T4T5T6T7T8T9T10Average
GT100.0±0.097.5±3.5100.0±0.095.0±0.075.0±7.125.0±0.097.5±3.5100.0±0.087.5±17.770.0±7.184.8±2.5
Direct Action97.5±3.5100.0±0.097.5±3.597.5±3.570.0±7.127.5±3.592.5±10.695.0±0.082.5±3.577.5±10.683.8±1.8
Spatial CBM*95.0±7.195.0±0.092.5±10.690.0±7.165.0±0.015.0±0.095.0±7.192.5±3.575.0±7.157.5±3.577.2±2.5
Spatial CBM90.0±0.085.0±0.095.0±0.090.0±7.162.5±3.512.5±3.595.0±0.087.5±10.675.0±7.157.5±3.575.0±0.7
Intervention
Spatial CBM, intervened95.0±0.087.5±3.597.5±3.595.0±0.067.5±3.517.5±3.597.5±3.595.0±0.082.5±10.662.5±3.579.8±2.5
Ablations
No concept90.0±7.180.0±7.187.5±3.585.0±0.062.5±3.515.0±0.090.0±7.172.5±3.562.5±3.537.5±3.568.2±2.5
GT, swapped†40.00.00.05.00.00.00.00.00.00.04.5

Success (%) on the ten LIBERO-Spatial tasks over two training seeds (†one seed).

Intervention

The key benefit of the bottleneck structure of CBMs is that they allow intervention of human-interpretable concepts at test time. They provide insight into the model's intent which it has to use to generate outputs. The intervener is provided all information of \(x_t\) and the predicted spatial concept \(\hat{c}_t\), and the oracle has additional access to the true concept \(c_t\). Based on the design of the spatial concept bottleneck, the intervener provides exocentric entities \(\tilde{a}^k_t\), while the robot's own ego entities are unchanged. We are thus able to re-generate the entire concept on intervention without changing all the exocentric entities.

\[ \tilde{c}^k_t = \psi\big(\tilde{a}^k_t, r\big), \qquad \tilde{y}_t = f(\tilde{c}_t, \rho_t) \]

Manipulation

In manipulation, the intervener provides the source or destination point \(\tilde{a}^k_t\), and the concept is regenerated. We use the oracle as the intervener and test three uses of the same operation. For correction, the oracle provides the true source and destination points at every step. For steering, we train the action expert with a single sentence shared across all ten tasks in place of the instruction, so that the concept is the only input that tells the robot which bowl to pick, and the oracle provides the position of the other, identical bowl as the source. For an unseen object, the oracle provides the position of the ramekin, which is never moved in any demonstration, and we compare against Direct Action whose instruction asks for the ramekin.

Manipulation interventions
Intervention on the manipulation concept, with two runs from the same start in each panel, without (top) and with (bottom) the intervention.

(a) Correction: the robot fails with the predicted concept \(\hat c\) and succeeds with the corrected \(\tilde c\).

(b) Steering: the intervener provides the other, identical bowl as the source \(\tilde a^1_t\), and the robot places that bowl instead.

(c) Unseen object: Direct Action does not place the ramekin when its instruction \(\tilde H\) asks for it, whereas providing the ramekin as the source \(\tilde a^1_t\) makes the robot place it.


Navigation

We split the spatial concept bottleneck of navigation, the costmap, into three sectors: left, center, and right, which map the three most common commands an intervener would likely give. The VLM is provided the last six camera frames taken five steps apart, eight frames from the earlier traversal, the current view, the sub-goal-containing view, and the final destination image. Most importantly, the VLM is provided only the predicted costmap \(\hat{c}_t\) (and never the true costmap \(c_t\)) to propose an intervention. It is required to output “none” if no change is required, or the proposed sector along with a brief reasoning. We use the oracle to detect differences between the predicted and true costmaps and trigger the intervention.

Each video shows one episode without (left) and with (right) intervention: the RGB view and the costmap (blue is low cost, white square: least-cost cell).

More qualitative results

Navigation: how to read these panels

Each panel shows one decision of the VLM. The top lines give the task and step, the VLM's chosen sector next to the oracle's and the predicted costmap's sectors, and the VLM's stated reason. The strip below shows the walkthrough frames the VLM was given, from the robot's position to the destination (green border). The left block is what the VLM saw, the current view with the predicted costmap over it, where the cross marks the costmap's lowest-cost cell. The middle block is the VLM's choice, where the circle marks the new target. The right column shows the reference and destination frames.

Navigation: successful interventions

Navigation: failure cases

The VLM's stated reason shows why. In Alt-goal the goal is not at the end of the recorded route, so relying on the route frames over the destination image misleads the VLM.

Manipulation

Each figure shows two runs from the same start: the top row without the intervention and the bottom row with it. Orange and blue boxes show where the concept places the source and the destination, and green marks the true positions.

Correction
Correction: the robot fails with the predicted concept and places the bowl with the corrected one.
Steering
Steering: with the other, identical bowl as the source, the robot places that bowl.
Unseen object
Unseen object: Direct Action does not place the ramekin, and with the ramekin as the source the robot places it on the plate.