Post-hoc Semantic Calibration for VLM-Human Preference Alignment
UrbanAlign is a training-free post-hoc calibration framework that aligns VLM outputs with human subjective preferences using minimal crowdsourced annotations. Through semantic dimension discovery, multi-agent deliberation, and hybrid-space calibration, it bridges the gap between VLM zero-shot judgments and human ground truth — without any model fine-tuning.
While demonstrated on urban scene perception (safety, beauty, liveliness, wealth, boringness, depressingness) using the Place Pulse 2.0 dataset, the framework is domain-agnostic: it applies to any pairwise preference task where VLMs can observe and humans can judge — food quality, interior design, landscape aesthetics, product appeal, and beyond. The core idea is general: extract interpretable evaluation dimensions from a few human-consensus examples, then use them to calibrate VLM scoring at scale.
Using TrueSkill-stratified consensus samples, the VLM discovers 5-8 universal evaluation dimensions (e.g., Facade Quality, Vegetation Coverage, Infrastructure Condition) with operational definitions — replacing hand-crafted rules with transferable, interpretable dimensions.
For each image pair, three VLM agents collaborate through deliberation:
- Observer — Describes visual evidence without premature judgment
- Debater — Argues opposing perspectives to reduce confirmation bias
- Judge — Synthesizes a final dimension-level score (1-10)
Constructs a hybrid embedding space (CLIP visual features + LLM semantic scores), then calibrates synthetic judgments against human ground truth via Locally Weighted Ridge Regression.
| Metric | Value |
|---|---|
| Average accuracy (6 categories) | 61.3% |
| Best single category (wealthy) | 72.3% |
| Gain over zero-shot VLM baseline | +17.9 pp |
| LWRR calibration gain | +7.0 pp |
| Cost reduction vs. full human annotation | 97% |
UrbanAlign/
├── urbanalign/ # Core package
│ ├── config.py # Global configuration (API, paths, hyperparams)
│ ├── preprocessing/
│ │ ├── extract_clip_features.py # [Optional] CLIP embedding extraction
│ │ ├── prepare_dataset.py # Filter & analyze Place Pulse annotations
│ │ ├── validate_data.py # Data quality checks
│ │ └── compute_trueskill.py # TrueSkill rating computation
│ ├── pipeline/
│ │ ├── stage1_semantic_extractor.py # Semantic dimension discovery
│ │ ├── stage2_multi_agent_synthesis.py # Multi-agent VLM scoring
│ │ └── stage3_hybrid_vrm.py # LWRR calibration & alignment
│ ├── evaluation/
│ │ ├── evaluate.py # Accuracy & kappa evaluation
│ │ ├── sensitivity_analysis.py # Hyperparameter sensitivity
│ │ ├── dimension_optimization.py # End-to-end dimension search
│ │ ├── results_summary.py # Result table generation
│ │ ├── exp3_self_consistency.py # Self-consistency (3×Mode2) ablation
│ │ ├── exp3_compute_lwrr.py # LWRR over self-consistency ensemble
│ │ ├── exp4_bootstrap_variance.py # Bootstrap CIs (full hyperparam search)
│ │ └── exp4b_bootstrap_fixed_params.py # Bootstrap CIs (fixed params)
│ └── baselines/
│ └── traditional_baselines.py # Siamese / Segmentation / Zero-shot baselines
├── scripts/
│ ├── run_all_modes.py # Run all Stage 2 modes sequentially
│ ├── specs_transfer_experiment.py # Cross-dataset transfer to SPECS
│ ├── specs_zeroshot.py # Zero-shot VLM baseline on SPECS
│ ├── ava_transfer_experiment.py # Cross-domain transfer to AVA (aesthetics)
│ ├── ava_baselines.py # Traditional baselines on AVA
│ └── generate_figures.py # Generate paper figures
├── assets/ # Framework diagrams
├── requirements.txt
├── .env.example
└── .gitignore
pip install -r requirements.txtFor CLIP feature extraction (optional — pre-extracted features can be reused):
pip install torch transformerscp .env.example .envFill in your .env:
URBANALIGN_API_KEY— API key for the VLM servicePLACE_PULSE_DIR— Path to the Place Pulse 2.0 datasetSPECS_DIR(optional) — Path to the SPECS dataset (for cross-dataset transfer)AVA_DIR(optional) — Path to the AVA dataset (for cross-domain aesthetic transfer)
Expected data layout:
<PLACE_PULSE_DIR>/
├── final_data_reliable_agg_N3.csv
├── final_data_reliable_raw_N3.csv
└── final_photo_dataset/
python -m urbanalign.preprocessing.extract_clip_features # Optional
python -m urbanalign.preprocessing.prepare_dataset
python -m urbanalign.preprocessing.compute_trueskillpython -m urbanalign.pipeline.stage1_semantic_extractor
python -m urbanalign.pipeline.stage2_multi_agent_synthesis
python -m urbanalign.pipeline.stage3_hybrid_vrmpython -m urbanalign.evaluation.evaluate
python -m urbanalign.evaluation.sensitivity_analysis
python -m urbanalign.baselines.traditional_baselines
python -m urbanalign.evaluation.results_summaryAll outputs are saved to urbanalign_outputs/ (auto-created, gitignored).
# Run all four Stage 2 modes sequentially, then evaluate
python -m scripts.run_all_modes
# Self-consistency ablation (3×Mode2 at equal API budget)
python -m urbanalign.evaluation.exp3_self_consistency
python -m urbanalign.evaluation.exp3_compute_lwrr
# Bootstrap confidence intervals
python -m urbanalign.evaluation.exp4_bootstrap_variance
python -m urbanalign.evaluation.exp4b_bootstrap_fixed_params
# Cross-dataset / cross-domain transfer (requires SPECS_DIR / AVA_DIR)
python -m scripts.specs_transfer_experiment
python -m scripts.specs_zeroshot
python -m scripts.ava_transfer_experiment
python -m scripts.ava_baselinesEdit urbanalign/config.py or set via environment variables:
| Parameter | Default | Description |
|---|---|---|
STAGE2_MODE |
4 | 1: single-direct, 2: pairwise-direct, 3: single-multiagent, 4: pairwise-multiagent |
N_POOL_MULTIPLIER |
0.01 | Sampling ratio from human annotation pool |
LABELED_SET_RATIO |
0.05 | Reference set ratio for LWRR calibration |
ALPHA_HYBRID |
0.3 | Hybrid weight (30% CLIP + 70% semantic) |
SELECTION_RATIO |
0.15 | Post-alignment retention ratio |
@article{zhang2026urbanalign,
title = {UrbanAlign: Post-hoc Semantic Calibration for
VLM-Human Preference Alignment},
author = {Zhang, Yecheng and others},
year = {2026}
}