All research notes

Model Development

Multimodal Research for the Messy Middle

Jul 5, 2026 · 6 min read

Benchmark multimodal tasks look like clean charts and well-lit photographs. Production multimodal tasks look like a fax of a form, photographed at an angle, partially handwritten, in two languages.

We build evaluation sets from that messy middle deliberately. Degradation curves under blur, skew, and compression tell us far more about deployment readiness than a leaderboard score.

Architecturally, the wins have come from resolution-adaptive encoding — spending vision tokens where the information density is, not uniformly across the page — and from explicit layout supervision rather than hoping the model infers structure.

The result is a model class that scores unremarkably on public benchmarks and dramatically better on the documents customers actually send us.