Breadth from the frontier
We use the best available LLMs for what they are good at: understanding messy input, reasoning across context, and producing a solid first pass. We don't train our own giant. We rent one.
Mitta Labs · AI product studio · United States
Mitta Labs builds vertical AI products for narrow, well-defined markets. We pair frontier large language models with small, task-specific models and hand-built domain tooling, so the output is finished work, not a first draft.
A 55 is impressive in a demo and unusable in production. Users don't pay for "mostly right". They pay for the last 25 points: the format that survives, the number that isn't hallucinated, the tone that fits the occasion.
Our thesis
General-purpose AI is a commodity. What is scarce is the engineering that turns a commodity into something a specific customer will pay for every month. Every Mitta product is built on the same three-part formula.
We use the best available LLMs for what they are good at: understanding messy input, reasoning across context, and producing a solid first pass. We don't train our own giant. We rent one.
Around the large model we place small, cheap, task-specific models and deterministic code: layout detection, number verification, style classifiers, quality gates. They do one thing, and they do it at 99%, not 55%.
We pick markets that are small enough to be ignored by platforms and painful enough to have a clear "finished" state. A memorial portrait. A notarized translation. A story a learner can actually read.
How we build
We have run this playbook four times. Each product took the same shape, which is why the fifth one will be faster than the fourth.
We write down what a finished output looks like for one specific customer, and what would make them reject it. That definition becomes the test suite.
Domain researchA frontier LLM produces the first pass: the translation, the story, the caption, the portrait direction. This is the 55.
Large modelSpecialised models and plain code fix what LLMs get wrong: they freeze numbers, preserve tables, classify tone, score readability, and reject anything off-spec.
Small models + codeNothing reaches a customer without passing automated quality gates. What fails becomes training data for the next small model.
Quality gatesMeasured on Chaptrio
Internal evaluation of Chaptrio's level-1 rewrites, before and after adding a deterministic readability ruler, blind cross-vendor judges and per-dimension score floors.
Why not just wait for a better model? Because the last 25 points are not a model problem. They are a product problem: knowing what "correct" means in one domain, and building the machinery to enforce it. That machinery compounds. A better LLM makes our products better on day one, and makes a prompt-wrapper competitor obsolete on the same day.
Products
Each one is small on purpose. Together they prove the formula transfers across languages, media types and customer segments.
Graded English reading for Chinese learners, built on the Focus on Form method. Nineteen classics, each rewritten at three levels with narration and illustrations. When a sentence stops a reader, one tap returns a precise diagnosis of what blocked them, then sends them straight back to the story. Read-aloud scoring, comprehension checks that can't be guessed, and a parent dashboard that shows the evidence.
Format-preserving translation for official documents: transcripts, certificates, bank statements, notarized paperwork. Tables stay tables. Numbers are frozen and verified by code, never by a model.
Self-service content production for creators on Xiaohongshu and WeChat Official Accounts. From a product and a positioning, it proposes topics, writes an internal brief, drafts a native post per platform, reviews it on five dimensions, and renders the image cards with accurate Chinese typography. A browser extension pastes the result into the editor. The creator presses publish.
Turns a single family photo into a dignified, print-ready memorial portrait for obituaries and services. Five styles, colour and black-and-white, 300 DPI output, preview before you pay.
Operating principles
If a customer can't tell you what finished looks like, we can't build quality gates for it. We walk away from vague markets, however large.
Frontier models are a utility, and we swap between them behind one interface. Our small models, readability rulers, evaluation sets and domain rules are the asset. They get better with every customer and don't depend on any single vendor.
Numbers, layouts, file formats and legal wording are verified by deterministic code. Prompts are for judgment, not for arithmetic.
A paying customer is the only reliable evaluator. Three of our four products launched with a price, and the fourth is invite-only until it earns one.
Shared infrastructure for generation, gating, billing and review means a new vertical is weeks of work, not quarters.
"Mitta" means measure in Finnish. Every product has a scored evaluation set, and we know our number before we ship.
For incubators & partners
Mitta Labs is a US company (Delaware) with a lean, technical founding team. We already have four products in market, three of them with live paid plans. What we're looking for from incubators and partners is not validation of the idea. It's leverage on distribution, compute and the next two verticals.
What we're looking for
Funeral homes, immigration and study-abroad agencies, language schools, creator networks.
Inference credits go directly into evaluation runs and small-model training.
We'd rather show live revenue and a live eval score than a pitch deck of projections.
Advice on when to double down on one product and when to spin up the next.
Get in touch
Investors, incubators, distribution partners and customers: one inbox, a human reply within two business days.