Theodoros Galanos
A working notebook on harness engineering, agentic systems, and the work of making AI useful in real engineering practice.
I work at the intersection of AI, design and engineering. The Harness is where I publish what I learn from building agentic systems and testing them against real engineering work.
My central question is what makes an agent's work trustworthy. That includes the tools it can use, the evidence it must produce, the checks around its outputs, and the decisions that remain with people.
Work you can inspect
- AEC-Bench: evaluating engineering agents. Experiments with HVAC tasks, model comparisons, and the effect of removing harness guidance.
- Auditing the benchmark itself. What looked like finished work and strong scores did not always survive engineering review. The published road-corridor run includes the source record and verifier receipt.
- Building recursive agents. The design and practical limits of Lambda-RLM on engineering tasks.
I also write about organisational authority and learning through repeated experiments. The technical and organisational conditions matter together.
For a broader view of my work, visit LinkedIn or GitHub. To follow new writing, subscribe by email or use the RSS feed.