Across 26 prompt conditions tested on Rust implementations (TLA+, Alloy, Verus, Kani, fuzzing), default unprompted coding outperformed most specialized verification prompts. Agents overwhelmingly wrote code first rather than modeling upfront: 75 out of 80 agents told to use TLA+ implemented in Rust first, then used TLA+ post-hoc on trivial functions where pass rates were already 99.6%, while completely failing to model concurrency, event queues, and state machines where the actual failures occurred. danluu.com