Aden Mann
Principal, Applied AI
Actuals · Australia
Aden Mann: Operator-grade
AI.
I build applied AI systems for work where being wrong is expensive. I'm Principal, Applied AI at Actuals; before that, four years at Immutable, a US$2.5B web3 company, building its AI and automation capability out of finance operations. Eleven years flying Army helicopters and an MBA came first.
The main theme for me is operations at the hard end: systems that break at bad moments, decisions with real stakes, and the gap between what a tool is supposed to do and what it actually does under pressure. Army aviation went digital partway through my career. Watching that transition up close, including unpacking where it went wrong as an aviation safety investigator, is where I developed a view on how human-machine work actually transfers, and how it fails.
01 Builds
Built and benchmarked. The numbers are attached.
AutoEvaluation
An optimisation engine that hill-climbs LLM instructions against a scoring function. It keeps what measurably works and reverts what doesn't, so a prompt improves itself while you sleep. Unlike DSPy, TextGrad and MIPRO, it needs no Python pipeline.
github.com/AdenCJM/AutoEvaluation ↗ (opens in a new tab)$ python3 autoeval.py --target instructions.md # each iteration: mutate → score → keep or revert judge deterministic 39/42 ✓ · agreement 0.91 ✓ result +25.2 points over baseline, no human in the loop
Parallel Research
One question, four models, answered at once and scored by agreement. Claude, GPT, Gemini and Perplexity run in parallel; a meta-pass marks where they converge and flags where a single model is out on its own.
github.com/AdenCJM/parallel-research ↗ (opens in a new tab)claude.md 7 findings · gpt.md 6 · gemini.md 8 · pplx.md 9 meta-analysis.md consensus 5 · unique 4 · contradiction 1 flagged ✓
AI Fluency Framework
A five-level, six-function matrix for measuring AI capability across a company. Built while rolling out AI at a scale-up, after seat count stopped being a useful metric: licences tell you who is paying, not who can actually use the thing.
Read the framework →02 Live agent
Running here, in production.
Ask AI Aden
An agent that answers the way I would, running live on this page.
“How do you know a prompt change actually made things better?”
03 Writing
Field notes on agents, evals and judgement.
- Speaker on applied AI and AI strategy, New York, Sydney and Melbourne
If you're working on something where being wrong is expensive, I'm interested in how you're approaching it.
aden@adenmann.com →