Topic

Aec Bench

7 articles in this topic.

  1. A World Worth Learning From

    Before an agent can learn from experience, its environment has to produce experience worth learning from. A sixteen-run engineering study tested what survives after the agent commits.

    aec-benchtask-worldsagent-evaluationreinforcement-learning
  2. The Attacker Moves Second. So Did I.

    My benchmark optimiser found the same seam an adversary would: it rewrote the world its own grader consumed. A design note on provenance ledgers—agent memory where authority comes from evidence, not persuasion.

    agent-securityagent-evaluationharness-engineeringaec-bench
  3. Fluent, But Unsafe

    How 150 supposedly finished tasks and perfect model scores hid a weak engineering benchmark—and how auditable reviews exposed what the numbers missed.

    agent-evaluationaec-benchtask-worldsai-in-aec
  4. Task Worlds and Meta-Harnesses

    How task worlds, Badiou, Plasticity, and the AEC-Bench meta-harness turn task prose, evidence, review, governance, and repair into runnable machinery.

    task-worldsharness-engineeringaec-benchagent-evaluation
  5. Plausible Answers, Failed Workflows

    An AEC-Bench release evaluation read as workflow reliability, not prose quality. Chapter by chapter: why a model can produce a plausible answer and still fail the durable record a project has to audit.

    agent-evaluationaec-benchai-in-aecharness-engineering
  6. Making aec-bench Trainable with Prime Lab

    How aec-bench and Prime Intellect's Lab turn engineering benchmarks into verifier-backed RL environments, adapter training runs, and inspectable traces.

    aec-benchreinforcement-learningagent-evaluationai-in-aec
  7. What HVAC Benchmarks Reveal About Agent Reliability

    Benchmarking AI agents on real HVAC engineering tasks across Claude and GPT models. Results on harness-dependent capability, agent evaluation design, and why AEC-domain benchmarks reveal what general benchmarks miss.

    harness-engineeringagent-evaluationai-in-aecaec-bench