Autonomous Driving Safety · 2026

Natural Language-Guided Generation of Long-Tail Critical Scenarios

Controllable hazard injection with constraint-preserving diffusion inpainting for closed-loop safety evaluation.

Tingting Lei · Yifan Zhu · Runxi Zhang · Feng Hu · Hong Yu · Ye Wang

Chongqing University of Posts and Telecommunications · Nanjing University of Aeronautics and Astronautics · Tongji University

01 · Abstract

Rare hazards need explicit control

Real-world driving logs undersample the interactions that matter most for safety. NLG-Gen converts language into executable spatial and kinematic constraints, writes the requested hazard into a structured scenario, and uses diffusion as a local realism restorer under semantic protection.

TL;DR. Explicit semantic editing establishes the long-tail hazard; mask-guided low-noise diffusion repairs the scene without erasing it; candidate filtering returns simulator-ready scenarios that challenge multiple planners.
02 · Method

Control first, restore second

The pipeline separates semantic correctness from generative realism instead of asking one soft conditioning signal to solve both.

Overview of the NLG-Gen pipeline

Natural-language intent parsing, vector-space hazard injection, constraint-preserving diffusion inpainting, and semantic candidate selection.

Language

Intent parsing

Produces scenario type, risk semantics, key agents, geometric relations, and severity.

Structure

Semantic editing

Injects explicit pedestrian-crossing, hard-brake, or cut-in conflicts in vector space.

Generation

Protected inpainting

Uses differential and semantic ROI masks with low-noise latent restoration.

Validation

Candidate filtering

Ranks candidates by semantic alignment, compliance, and preservation fidelity.

03 · Results

Controllable across hazards and severity

100%Compliance across all three hazard types
92.66Pedestrian-crossing MPA
94.42Aggressive-severity MPA
91.12%Repaint success rate
ScenarioMPA ↑CR ↑SQS ↑DRL ↑
Pedestrian crossing92.6610075.5135.98 m
Hard braking86.9310073.6440.09 m
Forced cut-in83.4810075.5743.57 m
04 · Analysis

Generated scenarios expose planner weaknesses

All three closed-loop planners degrade sharply on the generated set relative to the official nuPlan validation split.

PlannernuPlan score ↑NLG-Gen score ↑Collision-free ↑TTC ↑
PDM-Closed97.6127.8688.7576.50
Diffusion Planner95.7018.5969.7863.45
Flow Planner97.1310.0558.8151.86

What each stage contributes

  • Semantic editing raises MPA from 29.04 to 89.49.
  • Diffusion inpainting further raises MPA to 91.65 and interaction strength to 69.91.
  • The 41.03 CTC after repainting reveals the current realism-control trade-off.

Evaluation scope

  • Three hazard types and three severity levels.
  • nuPlan vector scenarios in a 64 m × 64 m ego-centric field.
  • Closed-loop tests with PDM-Closed, Diffusion Planner, and Flow Planner.
05 · Citation

Cite NLG-Gen

@article{lei2026nlggen,
  title  = {NLG-Gen: Natural Language-Guided Generation of Long-Tail Critical Scenarios for Autonomous Driving},
  author = {Lei, Tingting and Zhu, Yifan and Zhang, Runxi and Hu, Feng and Yu, Hong and Wang, Ye},
  year   = {2026}
}