Topic

Harness engineering

How tools, verifiers and control flow turn model capability into useful engineering work.

Start here

  1. Where Capability Actually Lives in Agentic Engineering

    In AEC and domain-specific engineering, AI agent capability lives not in the model alone but in harness engineering — the tools, verifiers, orchestration, and process design that make agentic work reliable.

  2. The Harness Is All You Need

    Why domain-specific agent harnesses, not bigger models, are what close the AI performance gap on real engineering tasks — and why the AEC industry needs proper benchmarks to prove it.

  3. Recursive by Design

    Building Recursive Language Model agents for real engineering tasks — from 1.5M tokens to 53K with Lambda-RLM, and what we learned about agent harness design along the way.

More in this topic

  1. Broad Creation, Narrow Authority

    An open-ended approach to AI-enabled software: let practitioners explore, embed controls in the platform, and govern the moment an experiment acquires organisational consequence.

    agentic-aiharness-engineeringgovernanceai-strategy
  2. The Attacker Moves Second. So Did I.

    My benchmark optimiser found the same seam an adversary would: it rewrote the world its own grader consumed. A design note on provenance ledgers—agent memory where authority comes from evidence, not persuasion.

    agent-securityagent-evaluationharness-engineeringaec-bench
  3. Mediation, Not Intermediation

    Why the 'fix your foundations before AI' message has it backwards: agentic workflows are the way out of legacy data, and governance worth having is co-designed from practice, not committees.

    harness-engineeringdata-engineeringgovernanceai-strategy
  4. Task Worlds and Meta-Harnesses

    How task worlds, Badiou, Plasticity, and the AEC-Bench meta-harness turn task prose, evidence, review, governance, and repair into runnable machinery.

    task-worldsharness-engineeringaec-benchagent-evaluation
  5. Plausible Answers, Failed Workflows

    An AEC-Bench release evaluation read as workflow reliability, not prose quality. Chapter by chapter: why a model can produce a plausible answer and still fail the durable record a project has to audit.

    agent-evaluationaec-benchai-in-aecharness-engineering
  6. Executable Standards

    Better tools and verifiers are not enough. The next harness boundary is the clause itself — turning standards, briefs, and codes into versioned predicates and replayable certificates.

    harness-engineeringformal-methodsai-in-aec
  7. The Third Axis

    What happens when you let the harness improve itself — two experiments in feedback-driven harness evolution, and an honest look at how rough the trajectory actually is.

    harness-engineeringautoresearchagentic-ai
  8. What If the Harness Could Improve Itself?

    Applying the autoresearch pattern to self-improve an engineering agent harness. Automated prompt optimisation across HVAC audit tasks on Claude and GPT-4.1-mini, showing how harness engineering compounds when the improvement loop runs itself.

    harness-engineeringautoresearchagent-evaluationai-in-aec
  9. What HVAC Benchmarks Reveal About Agent Reliability

    Benchmarking AI agents on real HVAC engineering tasks across Claude and GPT models. Results on harness-dependent capability, agent evaluation design, and why AEC-domain benchmarks reveal what general benchmarks miss.

    harness-engineeringagent-evaluationai-in-aecaec-bench