Illumination-Centric Diagnostic Benchmark

LumiBench

Benchmarking Illumination Robustness in Robotic Manipulation Policies

↗ Paper PDF ◎ Benchmark </> Code · Coming Soon

A controlled, mechanism-structured evaluation of how lighting changes alter robot-policy behavior.

LumiBench overview with illumination taxonomy, simulation conditions, severity control, and real-world evaluations
5illumination mechanisms
9simulated perturbations
5 × 10policies × RoboTwin tasks
RealARX physical validation
Motivation

Abstract

Illumination shifts can substantially degrade robotic manipulation policies, yet existing benchmarks typically treat lighting as a single perturbation factor. LumiBench instead organizes illumination variation by its dominant physical mechanism and evaluates it through paired nominal–perturbed rollouts.

For representative mechanisms, LumiBench adds within-condition severity sweeps to distinguish sensitivity to perturbation onset from degradation under increasing strength. The benchmark is instantiated in RoboTwin 2.0, evaluates five representative policy families across ten tasks, and is complemented by controlled physical experiments.

01

Mechanism-Structured

Intensity, chromaticity, spatial non-uniformity, temporal variation, and dynamic shadows.

02

Paired Evaluation

Task semantics, geometry, robot state, camera setup, checkpoint, and initialization are held fixed while illumination changes.

03

Severity + Physical Validation

Ordered severity sweeps expose response regimes, while controlled real-world trials test whether selected vulnerabilities persist.

Evaluation Space

Benchmark Design

LumiBench treats robustness as a controlled intervention problem: what changes when the illumination condition changes, and everything else is kept as constant as possible?

Intensity

Global illumination strength.

Dark

Chromaticity

Light-source color characteristics.

WarmCold

Spatial

Non-uniform illumination structure.

RandomSpotGrating

Temporal

Time-varying illumination.

Strobe

Shadow

Geometry-coupled moving shadows.

Moving Shadow
01

Broad illumination evaluation

One nominal Base condition plus nine controlled illumination perturbations across the five mechanisms. Disco Lighting serves as a composite stress condition.

02

Severity-controlled evaluation

Dark, Cold, Spot, and Strobe are swept over three ordered severity levels with physically interpretable controls.

03

Controlled physical evaluation

Selected illumination conditions are recreated on an ARX X5 platform to test whether simulation-identified sensitivity patterns persist.

Dark illumination condition
DarkIntensity
Warm illumination condition
WarmChromaticity
Cold illumination condition
ColdChromaticity
Random spatial lighting condition
RandomSpatial
Spot lighting condition
SpotSpatial
Grating lighting condition
GratingSpatial
Disco lighting condition
Disco LightingComposite
Strobe illumination condition
StrobeTemporal
Moving shadow condition
Moving ShadowShadow
Ordered Perturbation Strength

Severity-Controlled Evaluation

The benchmark separates the first cost of introducing a perturbation from additional degradation as that perturbation becomes stronger.

Mean success rate across five policies × ten tasks

Dark

L0 is Base. L1–L3 are ordered within each mechanism and are not assumed to be quantitatively comparable across mechanisms.

Per-policy success-rate changes under Dark, Cold, Spot and Strobe severity sweeps
Policy-specific severity response. The manuscript reports progressive, onset-dominated, and strongly policy-specific response patterns.
Evaluation

Illumination Robustness Is Not a Single Number

Different policies fail under different lighting mechanisms, and nominal task success does not reliably predict illumination robustness.

Illumination difficulty and policy-specific sensitivity heatmap
25.7 pp

Grating

Largest mean success-rate drop under the canonical benchmark settings.

24.0 pp

Disco

Composite chromatic, spatial, and temporal variation is also strongly disruptive.

18.5 pp

Spot

Localized spatial non-uniformity causes a large average drop and is reproduced in physical tests.

0.6 pp

Moving Shadow

Nearly harmless on average at the selected canonical setting, despite other lighting shifts being damaging.

Policy Summary

Base performance and mean retention

Mean retention is the average fraction of Base performance retained across the nine perturbations.

PolicyBase SR (%)Mean retention (%)Example sensitivity
RDT42.681.3Spatial + composite conditions
OpenVLA-OFT46.788.3Broad, relatively mild sensitivity
FastWAM54.346.1Large drops across many mechanisms
π0.573.990.4Mainly spatial non-uniformity
LingBot-VA48.779.0Temporal / composite sensitivity

These differences are observational benchmark results; the evaluation does not isolate architecture, pretraining data, or fine-tuning procedure as causal factors.

Qualitative Example

Lighting changes can alter closed-loop behavior

On Adjust Bottle, matched initial conditions under Base and Disco Lighting lead to visibly different end-effector trajectories, illustrating that illumination shift can propagate from perception into action.

End-effector trajectories under Base and Disco Lighting
Controlled Physical Evaluation

Selected Vulnerabilities Persist Under Real Lighting

Physical tests use an ARX X5 platform, fixed camera poses, disabled automatic exposure and white balance, and 16 trials per policy–task–condition configuration.

0/16

FastWAM · Adjust Bottle

After 16/16 Base success, all trials fail under Dark, Spot, and localized Strobe in the reported physical setup.

0%

π0.5 · Severe Spot

Severe Spot reduces all three evaluated π0.5 physical tasks to 0/16 success.

Strobe caveat

The physical Strobe uses intermittent switching of the same localized light as Spot, so it is a controlled stress condition—not a mechanism-matched realization of simulated global Strobe.

Physical Trials

Successful trials out of 16

PolicyTaskBaseWarmDark-MDark-SSpot-MSpot-SStrobe-MStrobe-S
π0.5Adjust Bottle16/1614/1614/1613/161/160/164/160/16
Beat Block Hammer14/1615/161/160/160/160/161/160/16
Stack Cube10/1612/168/160/160/160/160/160/16
FastWAMAdjust Bottle16/1614/160/160/160/160/160/160/16
Beat Block Hammer16/1616/166/167/160/160/160/160/16
Stack Cube0/160/160/160/160/160/160/160/16

Stack Cube for FastWAM is at 0/16 already under Base and therefore does not support a meaningful illumination-robustness comparison.

Reference

Citation

The current manuscript is anonymized, so the project page intentionally does not expose an author list or fabricate a public citation. Replace the citation block after the public version is released.

Public citation placeholder
LumiBench: Benchmarking Illumination Robustness in Robotic Manipulation Policies

Author names, venue status, arXiv identifier, repository URL, and dataset URL can be filled in from one small configuration area in index.html.

Enlarged project figure