Research

Frontier research. Production discipline.

Miniml’s leadership includes active academics — University of Edinburgh professors, an Alan Turing Institute Fellow, and an ELLIS Scholar. We work at the boundary of machine learning research and enterprise systems, then carry the methods that hold up under scrutiny into production. Research that ships.

Recent research from the team.

Technical work from Miniml researchers and collaborators — new methods, evaluations, and findings from frontier model development. Each entry links to a short summary on this site, with the full paper one click away.

Selected work — what it buys you
  • DeCoRe Why our retrieval systems don’t fabricate citations — decoding that suppresses hallucinated content before it reaches the user.
  • GRADA Hardening retrieval against poisoned documents — the checks we run before your RAG layer touches untrusted content.
  • MMLongBench When long context beats retrieval and when it doesn’t — measured, not guessed.
  • Q-Filters Cutting inference memory and serving cost without sacrificing accuracy in production.
  • Behavioral-shift auditing Knowing when a deployed model has quietly changed — the statistical test behind our drift monitoring.
  • FLARE Reasoning you can audit — traces that show why the system answered, not just what it answered.

How the research earns its place.

Most enterprise problems do not need a new method. They need known methods applied with discipline. Research enters our work where it genuinely moves the result — and the bar for that is high.

  • Evaluation first. Before a method ships, it is measured against your real data and real failure modes. The same evaluation rigor we use in research is how we decide what is ready for production.
  • Reliability over novelty. A dependable, well-understood approach beats a clever one that drifts. We reach for frontier methods only where a simpler approach genuinely cannot meet the reliability bar.
  • Novel methods where they pay. When a problem sits at the edge of what current tools can do, our academic leadership can bring methods from the literature — and the judgment to know when not to.
  • Built to own. What we ship is documented, testable, and handed over — so your team can run, extend, and re-evaluate it without us in the room.
Talk through a problem at the edge →
EVALUATION · RELIABILITY · GOVERNANCE Lab FRONTIER ML RESEARCH Method SELECTED · ADAPTED Evaluation REAL DATA · FAILURE MODES Hardening RELIABILITY · GUARDRAILS Production OWNED · GOVERNED
From the lab → method selected → evaluated on real data → hardened for reliability → production — with production behaviour fed back into evaluation
Start the conversation

Have a problem at the edge of what’s possible?

If your hardest problem needs more than off-the-shelf tools — or you simply want a second opinion grounded in current research — talk to us. A 30-minute conversation with a senior consultant who can tell you what is genuinely solvable today, and what it would take to ship it.

Book a consultation →