Agent Evaluation
7 pieces in this thread.
-
A World Worth Learning From
Before an agent can learn from experience, its environment has to produce experience worth learning from. A sixteen-run engineering study tested what survives after the agent commits.
-
The Attacker Moves Second. So Did I.
My benchmark optimiser found the same seam an adversary would: it rewrote the world its own grader consumed. A design note on provenance ledgers—agent memory where authority comes from evidence, not persuasion.
-
Fluent, But Unsafe
How 150 supposedly finished tasks and perfect model scores hid a weak engineering benchmark—and how auditable reviews exposed what the numbers missed.
-
Plausible Answers, Failed Workflows
An AEC-Bench release evaluation read as workflow reliability, not prose quality. Chapter by chapter: why a model can produce a plausible answer and still fail the durable record a project has to audit.
-
Making aec-bench Trainable with Prime Lab
How aec-bench and Prime Intellect's Lab turn engineering benchmarks into verifier-backed RL environments, adapter training runs, and inspectable traces.
-
What If the Harness Could Improve Itself?
Applying the autoresearch pattern to self-improve an engineering agent harness. Automated prompt optimisation across HVAC audit tasks on Claude and GPT-4.1-mini, showing how harness engineering compounds when the improvement loop runs itself.
-
Benchmarking Agents on Real Engineering Work Is Already Teaching Us Something Important
Benchmarking AI agents on real HVAC engineering tasks across Claude and GPT models. Results on harness-dependent capability, agent evaluation design, and why AEC-domain benchmarks reveal what general benchmarks miss.