AI in AEC
Bring agents into architecture, engineering and construction with room to experiment and clear limits on authority.
Start here
-
Broad Creation, Narrow Authority
An open-ended approach to AI-enabled software: let practitioners explore, embed controls in the platform, and govern the moment an experiment acquires organisational consequence.
-
Mediation, Not Intermediation
Why the 'fix your foundations before AI' message has it backwards: agentic workflows are the way out of legacy data, and governance worth having is co-designed from practice, not committees.
-
The Distance to the Edge
Why organisational AI capability is nonlinear and why the edge is better understood as a learning regime than a position on a maturity ladder.
More in this topic
-
Fluent, But Unsafe
How 150 supposedly finished tasks and perfect model scores hid a weak engineering benchmark—and how auditable reviews exposed what the numbers missed.
-
Plausible Answers, Failed Workflows
An AEC-Bench release evaluation read as workflow reliability, not prose quality. Chapter by chapter: why a model can produce a plausible answer and still fail the durable record a project has to audit.
-
Making aec-bench Trainable with Prime Lab
How aec-bench and Prime Intellect's Lab turn engineering benchmarks into verifier-backed RL environments, adapter training runs, and inspectable traces.
-
Executable Standards
Better tools and verifiers are not enough. The next harness boundary is the clause itself — turning standards, briefs, and codes into versioned predicates and replayable certificates.
-
Recursive by Design
Building Recursive Language Model agents for real engineering tasks — from 1.5M tokens to 53K with Lambda-RLM, and what we learned about agent harness design along the way.
-
The Harness Is All You Need
Why domain-specific agent harnesses, not bigger models, are what close the AI performance gap on real engineering tasks — and why the AEC industry needs proper benchmarks to prove it.
-
What If the Harness Could Improve Itself?
Applying the autoresearch pattern to self-improve an engineering agent harness. Automated prompt optimisation across HVAC audit tasks on Claude and GPT-4.1-mini, showing how harness engineering compounds when the improvement loop runs itself.
-
What HVAC Benchmarks Reveal About Agent Reliability
Benchmarking AI agents on real HVAC engineering tasks across Claude and GPT models. Results on harness-dependent capability, agent evaluation design, and why AEC-domain benchmarks reveal what general benchmarks miss.
-
Where Capability Actually Lives in Agentic Engineering
In AEC and domain-specific engineering, AI agent capability lives not in the model alone but in harness engineering — the tools, verifiers, orchestration, and process design that make agentic work reliable.