01

AI systems that work.
AI systems that you can trust: deployed, governed, verified.

AP Automation99% automation rate · 0% false positives
RAG Intelligence9.1/10 RAGAS · 200+ users · sub-3s latency
Agentic ERP90+ tools · human-in-the-loop · full AP/AR cycle
Fine-tuningdomain-specific · eval-driven · CI-gated

We build AI systems that write to live ERPs, code non-PO invoices automatically, and answer enterprise knowledge queries at 9.1/10 RAGAS accuracy — all governed by architecture, all evaluated in CI before they ship.

See what we buildView results →
99% automation rate
9.1 / 10 RAGAS accuracy
0% false positives
200+ active users
02

Four capabilities.
One engineering standard.

Every system Aria builds is evaluated against a golden dataset, traced end-to-end in Langfuse, and gated in CI before release.

03

Self-learning
invoice coding

Most AP automation requires professional services to configure and maintain coding rules. Ours learns from every confirmed invoice, compounding accuracy over time without manual intervention.

The system retrieves historically coded lines for the same vendor and cost centre, scores confidence, auto-posts high-confidence items, and routes edge cases to human review. Every confirmation writes back into the vector store.

Governance is proportional: high-confidence routine invoices are fully automated, novel or ambiguous items escalate with full audit trail. Zero false positives by design.

RAG retrievalConfidence gatingWrite-back loopQdrantOpenAI embeddingsFastAPILangfuse tracing
Automation rateoptimised corpus
99%
AccuracyGL code correctness
99%
False positive rateby architecture
0%
Recall@10retrieval precision
100%
Calibrationall confidence bands
100%
Corpus growthconfirmed lines
1,466→2,292
04
Active usersacross 2 tenants
200+
Retrieval accuracyRAGAS-scored
9.1/10
Extractabilityclear factual queries
9.4/10
Median latencyp50 response time
<3s
Vector chunkson Qdrant
50K+
First-contact resolutionconsultant queries
~60%
miles-rag
Production RAG pipeline — public reference implementation
↗

Knowledge assistants
that don't hallucinate

Generic RAG hallucinates on edge cases and retrieves the wrong chunk under load. Our assistants use hybrid BM25 + dense retrieval with weighted RRF — keyword precision where it matters, semantic understanding everywhere else.

Every release is gated in GitHub Actions against a 500+ question golden dataset. 9.0/10 minimum on RAGAS — any regression blocks the deploy. Custom LLM-as-judge evaluators score refusal behaviour and machine-consumable output separately.

Per-tenant Qdrant collection isolation, incremental S3 ingestion, three-tier Langfuse tracing. 15–20 senior consultant hours freed per week.

Hybrid retrievalBM25 + denseWeighted RRFQdrantRAGASLLM-as-judgeGitHub ActionsLangfuseMulti-tenant
05

Your ERP, in plain English.
Under governed, auditable control.

The ARIA middleware enforces approval queues, idempotency, and audit logging at the infrastructure layer — so governance is architectural, not prompt-based. No amount of prompt injection bypasses a mandatory approval queue.

90+ tools across the full AP/AR cycle. Invoices, receipts, payments, journals, credit notes — all accessible via Slack in natural language. New client onboarding reduced from multi-day to same-day.

—

Human-in-the-loop by architecture

Every write action routes through a mandatory approval queue with <2 minute average latency. Reads are instant. The distinction is enforced at the middleware layer, not in the prompt.

—

Semantic duplicate detection

Before any contact or account creation, a semantic MDM layer screens against existing records. Fuzzy name matching, address normalisation, cross-field similarity — catches duplicates that exact-match misses.

—

Three-tier distributed tracing

Every tool invocation is traced in Langfuse at message, tool execution, and downstream API call level. 50–100 daily tool invocations, fully observable. Defects found, fixed, and verified in a single session.

—

Trajectory evaluation in CI

Three judges — authorisation, action correctness, reversibility — over 103 audited production sessions. 0 gate bypasses, 0.0% hallucination, 100% judge-human agreement.

LangGraphFastAPISlack BoltPostgresRedisKubernetesSemantic MDMLangfuseHITL
06

Models trained on your domain,
evaluated before they ship.

General-purpose models underperform on domain-specific tasks — the vocabulary is wrong, the edge cases aren't represented, and there's no benchmark to measure improvement against. We fix that.

Domain-specific training on finance and enterprise workflows: GL coding, invoice classification, ERP entity extraction, document understanding. Every fine-tuned model is paired with a custom golden dataset and LLM-as-judge evaluators from day one.

Eval-driven development means the benchmark exists before the first training run. Improvements are measurable. Regressions are caught in CI.

Domain trainingEval-drivenLLM-as-judgeGolden datasetsCI gatingLangfuseRAGASCustom evaluators
01

Dataset design

Synthetic and real-world examples covering routine, near-edge, and hard-edge cases. Ground truth labelled before training begins.

02

Evaluator suite

Custom LLM-as-judge evaluators for your domain — not generic metrics. Scores extractability, correctness, and refusal behaviour separately.

03

Training runs

Domain-specific fine-tuning on Claude and GPT-family models. Iterative — each run measured against the same golden dataset.

04

CI gating

Every release gates on the golden dataset score. Minimum threshold enforced in GitHub Actions. No regression ships.

07

Six principles.
Applied to every system.

The stack changes. The methodology doesn't. Every system Aria builds is held to the same standard — regardless of which capability it delivers.

01

Governed by architecture

Approval queues, idempotency, and audit logging enforced at the infrastructure layer. Governance that holds under adversarial prompting.

02

Evaluated before it ships

Every release gates on a golden dataset in CI. RAGAS, custom LLM-as-judge evaluators, trajectory evaluation. If it regresses, it doesn't deploy.

03

Observable end-to-end

Three-tier Langfuse tracing at message, tool execution, and downstream API call level. Defects found, diagnosed, and fixed in a single session.

04

Self-learning by design

Write-back loops, corpus growth, and confidence recalibration. Every confirmed outcome improves the next prediction. Systems compound in value over time.

05

Context-engineered, not prompt-hacked

Retrieval, tool schemas, and multi-tenant grounding designed from first principles. The right context at inference — not a longer system prompt.

06

Built for production, not POC

Kubernetes, Postgres, Redis, S3, incremental ingestion, per-tenant isolation. The same engineering standards as the systems it integrates with.

08

Numbers from
production systems.

Not projections. Not benchmarks on toy datasets. These figures come from systems running in production, evaluated against golden datasets, traced in Langfuse.

0%
Automation rate
AP Automation POC
0%
Accuracy
GL code correctness
0%
False positives
by architecture
0/10
RAGAS score
9.1 avg ÷ 10
0+
Active users
across 2 tenants
0+
ERP tools
in production
0
Langfuse-scored confirmations
gate.correct = 1, every one
0
Audited ERP sessions
trajectory eval

Source: Langfuse export · langfuse_scores_export.csv · 1,213 rows · all gate.correct = 1

08

From 48-hour batch
to same-day automation.

A mid-market professional services firm was processing non-PO invoices manually — two finance analysts spending two days per week on GL coding, chasing approvals, and reconciling mispostings.

We deployed the AP Automation system against their existing invoice corpus. The RAG retrieval layer learned from 1,466 historically coded lines. Within the first week, 99% of routine invoices were coded and posted automatically — zero false positives, full audit trail.

The write-back loop grew the corpus to 2,292 confirmed lines by week four. The two analysts now review only genuinely novel or ambiguous invoices — roughly 1% of volume. The rest runs unattended.

Manual processing time
2 days / week< 2 hours / week−94%
Automation rate
0%99%+99pp
False positive rate
Unknown0%Eliminated
Corpus size
1,466 lines2,292 lines+56%
Onboarding time
Multi-day setupSame-day−95%

* Client anonymised. Figures from production system, Langfuse-traced.

09

Built for production.
Not a notebook.

Aria is Miles Waite — AI engineer and systems architect with 20 years of enterprise risk, trading systems, and real-time data infrastructure across European energy trading desks.

The systems on this site aren't prototypes. They serve 200+ active users, process hundreds of queries a day, and write to live ERPs under governance frameworks audited across 103 production sessions.

That background — risk architecture, real-time systems, production accountability — is why every Aria system is governed, observable, and evaluated. It's not a methodology. It's how you build when the stakes are real.

2022–
AI Engineer (Contract)
Form Consulting Services
2015–22
Senior Business Analyst
E.ON Energy Trading, Düsseldorf
2010–15
Lead BA, Enterprise Risk Platform
E.ON Energy Trading
2003–09
Sourcing Analyst
E.ON Energy Trading & HBOS
PythonFastAPILangChainLangGraphLangfuseQdrantRAGASNext.jsTypeScriptPostgresRedisKubernetesS3GitHub ActionsClaudeGPT-4
10

For AI systems that work
and you can trust —
contact:

Aria works with finance ops leaders and engineering teams at mid-market enterprises running modern ERPs. Contract or permanent, remote or hybrid.

LinkedIn →info@aria.ai