Messy sources
Public data comes in every format, shape, and form: zipped archives, PDFs, spreadsheets, binary grids, and feeds that were never meant for machines.
We turn messy public data like weather, economic releases, financial reports, or satellite imagery into reliable training datasets with live endpoints that power ML models in production.
THE PROBLEM
02 / 09Public data can drive predictions from energy and agriculture to climate and markets. Every team still has to acquire, join, and maintain the same sources first. That means months of engineering each time, and billions across the industry.
Public data comes in every format, shape, and form: zipped archives, PDFs, spreadsheets, binary grids, and feeds that were never meant for machines.
Every source uses different field names, units, and timestamps, so nothing joins until an engineer maps every field onto one shared schema.
The model needs the same data live, so the whole pipeline gets rebuilt a second time, and every source revision or outage silently breaks it.
THE SOLUTION
03 / 09Describe the prediction in your own words. Agents break it down and search our catalog, provider APIs, the open web, and even your own files. Every candidate is checked for coverage, cadence, and licence.
THE SOLUTION
04 / 09Agents write the code that ingests every source, cleans it, and joins it into one table your model can train on.
THE SOLUTION
05 / 09You watch every step happen in the app, from the search to the approval. Our cloud keeps the dataset live, and your model reads it through one endpoint.
Prague & Berlin daily maxima
question = {
'grain': 'One row per city and date',
'primary_key': 'city + date',
}
Column analytics
WHY WE BUILT IT
06 / 09We built this for ourselves first. Our prediction-market strategies trade on it. Every source we added cost real engineering time, and AI did not change that.
THE OPPORTUNITY
07 / 09Teams pay for the engineering and infrastructure to get public data into shape. They pay again for private data when public sources are not enough.
There are over 150,000 data engineers worldwide, and we estimate 15 to 20% of their time goes to public sources. This counts salary alone, so the real number is higher.
Our estimate · headcount from StartUs Insights, 2026Companies bought this much alternative data in 2025, from hedge funds and banks to retailers, insurers and logistics firms. We plan to offer these paid sources inside our search and take a cut of what teams buy.
Grand View and Fortune Business Insights, 2025 · six estimates span $12–25BGO TO MARKET + BUSINESS MODEL
08 / 09People already search for public datasets, and the results range from a clean API to a zip on an FTP. We take the most searched ones that are hardest to use, turn them into a proper API and publish them free.
Search by dataset, publisher, or topic
All datasetsClimateEconomyEnergyHousing
THE TEAM
09 / 09Vu and Vojta have been friends for 20 years, since they built a mobile game together in high school. Vu met Rob at Marinade. The three of us started Mostly Right in early 2026.

Shipped 13 developer products. Built Avocode through 500 Startups to $2M ARR before it was acquired.

Founded Cinnamon and raised a $3.9M seed. Led product at Marinade Finance for two years.

Masters at Imperial, quantum computing PhD at Oxford, now at IBM. His first-author Nature paper has almost 4,000 citations.