01 / 09

Data infrastructure for real-world predictions.

We turn messy public data like weather, economic releases, financial reports, or satellite imagery into reliable training datasets with live endpoints that power ML models in production.

02 / 09

Public data is free. Using it is not.

Public data can drive predictions from energy and agriculture to climate and markets. Every team still has to acquire, join, and maintain the same sources first. That means months of engineering each time, and billions across the industry.

01

Messy sources

Public data comes in every format, shape, and form: zipped archives, PDFs, spreadsheets, binary grids, and feeds that were never meant for machines.

02

Nothing lines up

Every source uses different field names, units, and timestamps, so nothing joins until an engineer maps every field onto one shared schema.

03

Stale in production

The model needs the same data live, so the whole pipeline gets rebuilt a second time, and every source revision or outage silently breaks it.

03 / 09

Our agents find and check the sources.

Describe the prediction in your own words. Agents break it down and search our catalog, provider APIs, the open web, and even your own files. Every candidate is checked for coverage, cadence, and licence.

04 / 09

Then they build and deploy the dataset.

Agents write the code that ingests every source, cleans it, and joins it into one table your model can train on.

Zero ops
You approve the dataset once. It deploys to our infrastructure and keeps itself updated.
One clean API
Whatever the sources were, you get the latest rows from one endpoint, in the same shape every time.
05 / 09

One place to put public data into production.

You watch every step happen in the app, from the search to the approval. Our cloud keeps the dataset live, and your model reads it through one endpoint.

06 / 09

We know this is a real problem.

We built this for ourselves first. Our prediction-market strategies trade on it. Every source we added cost real engineering time, and AI did not change that.

Ten teams on the SDK
We exposed it as an SDK that puts live and historical data behind one schema. Ten teams use it and one pays us.
Then we built the harness
Every team asks for different sources. Writing each adapter by hand does not scale, so the harness does it instead.
07 / 09

And it’s not small. Teams already spend billions on data.

Teams pay for the engineering and infrastructure to get public data into shape. They pay again for private data when public sources are not enough.

Spent on engineering $3–6B

There are over 150,000 data engineers worldwide, and we estimate 15 to 20% of their time goes to public sources. This counts salary alone, so the real number is higher.

Our estimate · headcount from StartUs Insights, 2026
Spent on private data $19B

Companies bought this much alternative data in 2025, from hedge funds and banks to retailers, insurers and logistics firms. We plan to offer these paid sources inside our search and take a cut of what teams buy.

Grand View and Fortune Business Insights, 2025 · six estimates span $12–25B
08 / 09

Public data is free. We get paid to run it.

People already search for public datasets, and the results range from a clean API to a zip on an FTP. We take the most searched ones that are hardest to use, turn them into a proper API and publish them free.

Free, so they find us
Anyone looking for that data lands on our version. They can use it without paying us anything.
Paid when they fork it
Teams putting a model in production soon need another source or a different shape. They run that dataset on our infrastructure and pay for compute, storage and bandwidth.
09 / 09

Founders who have shipped, sold and published.

Vu and Vojta have been friends for 20 years, since they built a mobile game together in high school. Vu met Rob at Marinade. The three of us started Mostly Right in early 2026.

Vu Hoang Anh

Vu Hoang Anh

CEO

Shipped 13 developer products. Built Avocode through 500 Startups to $2M ARR before it was acquired.

Rob Tarabcak

Rob Tarabcak

CTO

Founded Cinnamon and raised a $3.9M seed. Led product at Marinade Finance for two years.

Vojta

Vojta Havlicek

RESEARCH

Masters at Imperial, quantum computing PhD at Oxford, now at IBM. His first-author Nature paper has almost 4,000 citations.

Backed by Hustle Fund 2080 Ventures