Topic

Agent evaluation

Start with measured results, then examine what scores miss and what a trustworthy task needs.

Start here

  1. What HVAC Benchmarks Reveal About Agent Reliability

    Benchmarking AI agents on real HVAC engineering tasks across Claude and GPT models. Results on harness-dependent capability, agent evaluation design, and why AEC-domain benchmarks reveal what general benchmarks miss.

  2. Fluent, But Unsafe

    How 150 supposedly finished tasks and perfect model scores hid a weak engineering benchmark—and how auditable reviews exposed what the numbers missed.

  3. A World Worth Learning From

    Before an agent can learn from experience, its environment has to produce experience worth learning from. A sixteen-run engineering study tested what survives after the agent commits.

More in this topic

  1. The Attacker Moves Second. So Did I.

    My benchmark optimiser found the same seam an adversary would: it rewrote the world its own grader consumed. A design note on provenance ledgers—agent memory where authority comes from evidence, not persuasion.

    agent-securityagent-evaluationharness-engineeringaec-bench
  2. Task Worlds and Meta-Harnesses

    How task worlds, Badiou, Plasticity, and the AEC-Bench meta-harness turn task prose, evidence, review, governance, and repair into runnable machinery.

    task-worldsharness-engineeringaec-benchagent-evaluation
  3. Plausible Answers, Failed Workflows

    An AEC-Bench release evaluation read as workflow reliability, not prose quality. Chapter by chapter: why a model can produce a plausible answer and still fail the durable record a project has to audit.

    agent-evaluationaec-benchai-in-aecharness-engineering
  4. Making aec-bench Trainable with Prime Lab

    How aec-bench and Prime Intellect's Lab turn engineering benchmarks into verifier-backed RL environments, adapter training runs, and inspectable traces.

    aec-benchreinforcement-learningagent-evaluationai-in-aec
  5. What If the Harness Could Improve Itself?

    Applying the autoresearch pattern to self-improve an engineering agent harness. Automated prompt optimisation across HVAC audit tasks on Claude and GPT-4.1-mini, showing how harness engineering compounds when the improvement loop runs itself.

    harness-engineeringautoresearchagent-evaluationai-in-aec