SABER: Learning Attention-based Semantic Affordance for Legged Locomotion

Hari Prasanth Palanivelu, Samuel Sze, Kennard Garrison Johannes, Albertus Hendrawan Adiwahono, Meng Yee (Michael) Chuah
Institute of Advanced Intelligence and Computing (IAIC), Agency for Science, Technology and Research (A*STAR), Singapore
arXiv PDF Video BibTeX

Abstract

Perceptive legged locomotion has advanced rapidly by integrating terrain geometry into learned policies, yet the integration of terrain meaning remains sparse: a pipe, a patch of grass, or a fragile box may be geometrically traversable while being inappropriate for contact. In industrial environments, where legged robots increasingly operate, a single misplaced step can damage fragile equipment, destabilize the robot, or endanger the site. To address this, we introduce SABER, a planner-free reinforcement-learning policy that jointly reasons about terrain geometry and semantic contact permission. The policy consumes a unified terrain-affordance map, where each cell encodes local 3D geometry and a semantic contact cost. We augment cross-attention with a learned, signed semantic bias: an additive term on the attention logits, gated by the contact cost, that reweights flagged cells by their distance from the nearest foot. A hazard therefore reshapes attention where it can still affect the next foothold, and its influence fades where it cannot. The resulting policy selects footholds on permitted support and keeps the leg clear of forbidden regions throughout the swing phase. We perform a systematic ablation that isolates the contribution of each architectural component; removing the semantic bias alone increases forbidden contacts by 55% while velocity tracking is unchanged. We validate the policy on a Unitree B2, demonstrating sim-to-real semantic contact selection across indoor and outdoor environments and four semantic obstacle classes.

Hardware results

Semantic affordance on the Unitree B2

Perception, mapping and policy inference all run onboard the robot, and velocity commands come from a wireless joystick.

Click any video to zoom.

videos/indoor_stepping_stones.mp4
Indoor

Semantic stepping stones

videos/indoor_pipe_30cm.mp4
Indoor

30 cm pipe

videos/indoor_turf_sideways.mp4
Indoor

Turf patches

videos/indoor_three_low_pipes.mp4
Indoor

Three low pipes, laid parallel

videos/outdoor_pipe_deck.mp4
Outdoor

Outdoor pipe

videos/outdoor_grass_strip.mp4
Outdoor

Grass strip beside a pavement

Controlled comparison: same policy, segmentation on vs off

The outdoor tiles with grass gaps are the hardware counterpart of the lava-tile terrain. Both runs use the identical checkpoint; only the cost channel changes.

videos/outdoor_tiles_without_semantics.mp4
Without semantics
videos/outdoor_tiles_with_semantics.mp4
With semantics
Method

SABER Architecture

SABER architecture: a CNN tokenizes the 41x21x4 terrain-affordance map into 231 tokens; a proprioceptive query attends over them with eight heads; the contact cost enters as a CNN channel and as a distance-structured semantic bias on the attention logits; an MLP outputs joint targets for the Unitree B2
A CNN tokenizes the 41×21×4 terrain-affordance map (x, y, z, r) into 231 terrain tokens; a single proprioceptive query attends over them across eight heads. The contact cost enters both as a CNN input channel and as a distance-structured additive bias on the attention logits.

The semantic attention bias

Each head h adds a bias to the logit of terrain token i:

Bh,i = βh · ρh(di) · r̄i

βh is a signed per-head gain, ρh a learned radial profile over the distance di from the cell to its nearest foot, and r̄i the max-pooled contact cost. The bias is not a mask: flagged cells still receive attention, but its weight is reshaped where it can still change the next foothold and fades where it cannot. Eight gains and eight six-point profiles are the whole mechanism, so the learned behaviour is directly inspectable.

Learned semantic attention bias: effective bias per head versus distance to the nearest foot, with most heads suppressing flagged terrain, one emphasising it near the foot, and all decaying to zero beyond about 1 m; the bias shifts roughly three percentage points of attention mass away from flagged cells on every terrain
(a) Effective bias of each head vs. distance to the nearest foot: one head emphasises flagged terrain near the foot, most suppress it, and all decay toward zero beyond ~1 m. (b) Attention mass shifted away from flagged cells per terrain.

Visualization of learnt attention

Head-averaged attention over the 231 terrain tokens on three training terrains. Sphere size and colour encode attention weight (uniform would be 1/231 ≈ 0.004). Attention collapses onto the landing cell toward touchdown and leaves almost none on flagged terrain at the moment that decides where the foot goes.

Concentric grass rings
Pipes
Semantic stepping stones
Simulation

Training terrains

(a) Rough terrain; (b) pipes; (c) semantic stepping stones; (d) rails; (e) lava tiles; (f) concentric grass rings. Red points are flagged cells, blue are unflagged.

Architecture ablation

Six variants, each changing one element with everything else fixed. Averaged over five evaluation terrains, 256 robots each, ten seeds. Contacts are flagged-terrain contact events per episode; velocity error in m/s; distance in m.

VariantContacts ↓Vel. err. ↓Success ↑Dist. ↑
SABER1.040.1850.9086.15
SABER, 16 heads1.110.1730.8976.23
SABER, 32 heads1.920.1990.8605.96
Static bias, w/o uℓ1.710.2210.8965.94
w/o bias1.610.1850.9056.18
w/o bias, w/o uℓ1.700.2300.9055.89
w/o semantics, w/o bias, w/o uℓ13.510.1780.9056.11
Bar chart of flagged-terrain foot contacts per episode on semantic stepping stones, concentric grass rings and lava tiles: SABER 0.68, 1.52 and 1.34 versus 1.26, 2.33 and 2.29 without the semantic bias

Removing the semantic bias alone raises flagged contacts by 55% while velocity error, success and distance are unchanged. Removing the cost channel entirely raises contacts thirteen-fold: the policy simply walks through flagged regions.

Flagged-terrain contacts on the three terrains where flagged and unflagged ground are geometrically identical. The gain is largest exactly where geometry offers no cue.

Citation

BibTeX

@misc{palanivelu2026saber,
  title         = {{SABER}: Learning Attention-based Semantic Affordance for Legged Locomotion},
  author        = {Palanivelu, Hari Prasanth and Sze, Samuel and Johannes, Kennard Garrison and Adiwahono, Albertus Hendrawan and Chuah, Meng Yee (Michael)},
  year          = {2026},
  eprint        = {2609.21572},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  url           = {https://arxiv.org/abs/2609.21572}
}