AI Engineering · Enterprise Finance
AI systems that work.
AI systems that you can trust: deployed, governed, verified.
We build AI systems that write to live ERPs, code non-PO invoices automatically, and answer enterprise knowledge queries at 9.1/10 RAGAS accuracy — all governed by architecture, all evaluated in CI before they ship.
Capabilities
Four capabilities.
One engineering standard.
Every system Aria builds is evaluated against a golden dataset, traced end-to-end in Langfuse, and gated in CI before release.
AP Automation
Self-learning
invoice coding
Most AP automation requires professional services to configure and maintain coding rules. Ours learns from every confirmed invoice, compounding accuracy over time without manual intervention.
The system retrieves historically coded lines for the same vendor and cost centre, scores confidence, auto-posts high-confidence items, and routes edge cases to human review. Every confirmation writes back into the vector store.
Governance is proportional: high-confidence routine invoices are fully automated, novel or ambiguous items escalate with full audit trail. Zero false positives by design.
RAG Intelligence
Knowledge assistants
that don't hallucinate
Generic RAG hallucinates on edge cases and retrieves the wrong chunk under load. Our assistants use hybrid BM25 + dense retrieval with weighted RRF — keyword precision where it matters, semantic understanding everywhere else.
Every release is gated in GitHub Actions against a 500+ question golden dataset. 9.0/10 minimum on RAGAS — any regression blocks the deploy. Custom LLM-as-judge evaluators score refusal behaviour and machine-consumable output separately.
Per-tenant Qdrant collection isolation, incremental S3 ingestion, three-tier Langfuse tracing. 15–20 senior consultant hours freed per week.
Agentic ERP Automation
Your ERP, in plain English.
Under governed, auditable control.
The ARIA middleware enforces approval queues, idempotency, and audit logging at the infrastructure layer — so governance is architectural, not prompt-based. No amount of prompt injection bypasses a mandatory approval queue.
90+ tools across the full AP/AR cycle. Invoices, receipts, payments, journals, credit notes — all accessible via Slack in natural language. New client onboarding reduced from multi-day to same-day.
Human-in-the-loop by architecture
Every write action routes through a mandatory approval queue with <2 minute average latency. Reads are instant. The distinction is enforced at the middleware layer, not in the prompt.
Semantic duplicate detection
Before any contact or account creation, a semantic MDM layer screens against existing records. Fuzzy name matching, address normalisation, cross-field similarity — catches duplicates that exact-match misses.
Three-tier distributed tracing
Every tool invocation is traced in Langfuse at message, tool execution, and downstream API call level. 50–100 daily tool invocations, fully observable. Defects found, fixed, and verified in a single session.
Trajectory evaluation in CI
Three judges — authorisation, action correctness, reversibility — over 103 audited production sessions. 0 gate bypasses, 0.0% hallucination, 100% judge-human agreement.
Model Fine-tuning
Models trained on your domain,
evaluated before they ship.
General-purpose models underperform on domain-specific tasks — the vocabulary is wrong, the edge cases aren't represented, and there's no benchmark to measure improvement against. We fix that.
Domain-specific training on finance and enterprise workflows: GL coding, invoice classification, ERP entity extraction, document understanding. Every fine-tuned model is paired with a custom golden dataset and LLM-as-judge evaluators from day one.
Eval-driven development means the benchmark exists before the first training run. Improvements are measurable. Regressions are caught in CI.
Dataset design
Synthetic and real-world examples covering routine, near-edge, and hard-edge cases. Ground truth labelled before training begins.
Evaluator suite
Custom LLM-as-judge evaluators for your domain — not generic metrics. Scores extractability, correctness, and refusal behaviour separately.
Training runs
Domain-specific fine-tuning on Claude and GPT-family models. Iterative — each run measured against the same golden dataset.
CI gating
Every release gates on the golden dataset score. Minimum threshold enforced in GitHub Actions. No regression ships.
How we build
Six principles.
Applied to every system.
The stack changes. The methodology doesn't. Every system Aria builds is held to the same standard — regardless of which capability it delivers.
Governed by architecture
Approval queues, idempotency, and audit logging enforced at the infrastructure layer. Governance that holds under adversarial prompting.
Evaluated before it ships
Every release gates on a golden dataset in CI. RAGAS, custom LLM-as-judge evaluators, trajectory evaluation. If it regresses, it doesn't deploy.
Observable end-to-end
Three-tier Langfuse tracing at message, tool execution, and downstream API call level. Defects found, diagnosed, and fixed in a single session.
Self-learning by design
Write-back loops, corpus growth, and confidence recalibration. Every confirmed outcome improves the next prediction. Systems compound in value over time.
Context-engineered, not prompt-hacked
Retrieval, tool schemas, and multi-tenant grounding designed from first principles. The right context at inference — not a longer system prompt.
Built for production, not POC
Kubernetes, Postgres, Redis, S3, incremental ingestion, per-tenant isolation. The same engineering standards as the systems it integrates with.
Results
Numbers from
production systems.
Not projections. Not benchmarks on toy datasets. These figures come from systems running in production, evaluated against golden datasets, traced in Langfuse.
Source: Langfuse export · langfuse_scores_export.csv · 1,213 rows · all gate.correct = 1
Case study
From 48-hour batch
to same-day automation.
A mid-market professional services firm was processing non-PO invoices manually — two finance analysts spending two days per week on GL coding, chasing approvals, and reconciling mispostings.
We deployed the AP Automation system against their existing invoice corpus. The RAG retrieval layer learned from 1,466 historically coded lines. Within the first week, 99% of routine invoices were coded and posted automatically — zero false positives, full audit trail.
The write-back loop grew the corpus to 2,292 confirmed lines by week four. The two analysts now review only genuinely novel or ambiguous invoices — roughly 1% of volume. The rest runs unattended.
* Client anonymised. Figures from production system, Langfuse-traced.
About
Built for production.
Not a notebook.
Aria is Miles Waite — AI engineer and systems architect with 20 years of enterprise risk, trading systems, and real-time data infrastructure across European energy trading desks.
The systems on this site aren't prototypes. They serve 200+ active users, process hundreds of queries a day, and write to live ERPs under governance frameworks audited across 103 production sessions.
That background — risk architecture, real-time systems, production accountability — is why every Aria system is governed, observable, and evaluated. It's not a methodology. It's how you build when the stakes are real.
Background
Stack
Get in touch
For AI systems that work
and you can trust —
contact:
Aria works with finance ops leaders and engineering teams at mid-market enterprises running modern ERPs. Contract or permanent, remote or hybrid.