Skip to content

Add Provael under Data, Simulation & Benchmarks - #1

Open
sattyamjjain wants to merge 1 commit into
ai4s-research:mainfrom
sattyamjjain:add-provael
Open

sattyamjjain wants to merge 1 commit into
ai4s-research:mainfrom
sattyamjjain:add-provael

Conversation

@sattyamjjain

Copy link
Copy Markdown

Adds one entry to 🧪 Data, Simulation & Benchmarks, in the section's format: - [Title](URL) — note (year).

Disclosure: I am the author.

The section covers benchmarks that measure capability — LIBERO, SIMPLER, CALVIN, RoboCasa, VLABench. This is the adversarial counterpart: it perturbs the instruction and observation a policy receives inside LIBERO or Meta-World and asks whether the end-effector left a keep-out region, reporting the rate with a task-clustered 95% CI against a matched benign control. The list currently has no safety or security entry at all, which is why I thought it was worth proposing rather than a stretch.

Software, not a paper. Every other row links an arXiv id; this has none, and the entry says so rather than implying a publication. If a paperless row does not fit the section, closing this is the right call.

Limits: simulation only, one policy measured on one suite, low single-digit stars, two months old. Two of its three measured attack families scored 0% against the real model and that is published at the same prominence as the positive result.

Link verified resolving.

The section covers capability benchmarks; this is the adversarial counterpart.
Software rather than a paper, stated in the entry since every other row is an
arXiv link. Format follows CONTRIBUTING: - [Title](URL) - note (year).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant