Archive

Writing

Notes on harness engineering, agent evaluation, and building reliable AI systems for real engineering work.

2026 16 pieces
  1. The Distance to the Edge

    Why organisational AI capability is nonlinear and why the edge is better understood as a learning regime than a position on a maturity ladder.

    ai-strategyorganisational-learningfrontier-aiai-capability
  2. A World Worth Learning From

    Before an agent can learn from experience, its environment has to produce experience worth learning from. A sixteen-run engineering study tested what survives after the agent commits.

    harness-engineeringagent-evaluationaec-benchtask-worlds
  3. Broad Creation, Narrow Authority

    An open-ended approach to AI-enabled software: let practitioners explore, embed controls in the platform, and govern the moment an experiment acquires organisational consequence.

    agentic-aiharness-engineeringgovernanceai-strategy
  4. The Attacker Moves Second. So Did I.

    My benchmark optimiser found the same seam an adversary would: it rewrote the world its own grader consumed. A design note on provenance ledgers—agent memory where authority comes from evidence, not persuasion.

    harness-engineeringagent-evaluationaec-benchagent-security
  5. Fluent, But Unsafe

    How 150 supposedly finished tasks and perfect model scores hid a weak engineering benchmark—and how auditable reviews exposed what the numbers missed.

    harness-engineeringagent-evaluationaec-benchtask-worlds
  6. Mediation, Not Intermediation

    Why the 'fix your foundations before AI' message has it backwards: agentic workflows are the way out of legacy data, and governance worth having is co-designed from practice, not committees.

    agentic-aiharness-engineeringdata-engineeringgovernance
  7. Task Worlds and Meta-Harnesses

    How task worlds, Badiou, Plasticity, and the AEC-Bench meta-harness turn task prose, evidence, review, governance, and repair into runnable machinery.

    agentic-aiharness-engineeringaec-benchtask-worlds
  8. Plausible Answers, Failed Workflows

    An AEC-Bench release evaluation read as workflow reliability, not prose quality. Chapter by chapter: why a model can produce a plausible answer and still fail the durable record a project has to audit.

    harness-engineeringagentic-aiai-in-aecai-benchmarks
  9. Making aec-bench Trainable with Prime Lab

    How aec-bench and Prime Intellect's Lab turn engineering benchmarks into verifier-backed RL environments, adapter training runs, and inspectable traces.

    aec-benchprime-labreinforcement-learningagent-evaluation
  10. Executable Standards

    Better tools and verifiers are not enough. The next harness boundary is the clause itself — turning standards, briefs, and codes into versioned predicates and replayable certificates.

    harness-engineeringautoformalisationai-in-aecformal-methods
  11. The Third Axis

    What happens when you let the harness improve itself — two experiments in feedback-driven harness evolution, and an honest look at how rough the trajectory actually is.

    harness-engineeringagentic-aiself-improvementautoresearch
  12. What If the Harness Could Improve Itself?

    Applying the autoresearch pattern to self-improve an engineering agent harness. Automated prompt optimisation across HVAC audit tasks on Claude and GPT-4.1-mini, showing how harness engineering compounds when the improvement loop runs itself.

    harness-engineeringautoresearchagentic-aiai-in-aec
  13. Benchmarking Agents on Real Engineering Work Is Already Teaching Us Something Important

    Benchmarking AI agents on real HVAC engineering tasks across Claude and GPT models. Results on harness-dependent capability, agent evaluation design, and why AEC-domain benchmarks reveal what general benchmarks miss.

    harness-engineeringagentic-aiai-in-aecai-benchmarks
  14. Where Capability Actually Lives in Agentic Engineering

    In AEC and domain-specific engineering, AI agent capability lives not in the model alone but in harness engineering — the tools, verifiers, orchestration, and process design that make agentic work reliable.

    harness-engineeringagentic-aiai-in-aecengineering-ai