From messy sources to one clean, live API.

Let agents find the sources, clean and join the data, and keep it updated. Download the full Parquet for training or query the latest rows through one API call.

Projects Climate research Datasets
DatasetVersionShapeVisibility
  • Prague & Berlin daily maximaParquet · versioned · refreshed 7 minutes ago v18 12,162 rows × 8 columns Public
  • Central Europe rainfallParquet · versioned · refreshed yesterday v2 18,204 rows × 6 columns Private
  • Alpine station coverageParquet · versioned · refreshed Aug 16 v4 2,418 rows × 7 columns Public
  • Vltava basin observationsParquet · versioned · refreshed Aug 15 v3 91,302 rows × 7 columns Private
  • Central Europe heat alertsParquet · versioned · refreshed Aug 13 v5 8,741 rows × 11 columns Link only
  • Prague hourly weatherParquet · versioned · refreshed Aug 11 v8 142,560 rows × 17 columns Private
  • Berlin station inventoryParquet · versioned · refreshed Aug 8 v2 1,084 rows × 22 columns Private
  • Vltava flood eventsParquet · versioned · refreshed Aug 5 v3 628 rows × 31 columns Public
  • Regional climate normalsParquet · versioned · refreshed Jul 28 v6 54,216 rows × 19 columns Public
  • Station elevation bandsParquet · versioned · refreshed Jul 23 v1 2,419 rows × 9 columns Private

Daily maximum temperature for Prague and Berlin from 2010 to today.

The question

question = {
    'grain': 'One row per city and date',
    'primary_key': 'city + date',
}
One row
One row per city and calendar date
Required fields
city, date, max_temp_c, station_id, country, source, observed_at, ingested_at
Feasibility
supportable

Sources

  • GHCN-HourlyNOAA · our catalogmatched
  • ASOS station feedIowa Environmental Mesonet · provider APImatched
  • DWD open dataDeutscher Wetterdienst · open webmatched

41 candidates checked · 3 cover both cities for the whole range

Build

Joined
3 sources on city + date
Normalized
°F and kelvin to °C, local time to UTC
Result
12,162 rows × 8 columns

Column analytics

max_temp_cfloat64 · °C
Missing
0
Distinct
624
Median
13.7 °C
−18°−5°21°34°47°

Deployed

GET/v1/prague-berlin-daily-maxima

Status
Live, no pipeline to run
Refresh
Hourly, and on every upstream revision
Version
v18 · schema pinned

From a data question to a dataset you can use.

Start with a data question. The result is a clean dataset that stays current and is ready to download or query.

Find the rightsources.

Agents search datasets and public sources, then compare coverage, history, fields, and update cycles.

Read everyformat.

Agents connect to open and keyed APIs, then read data from CSV, JSON, ZIP, PDF, and GRIB files.

Clean and jointhe data.

Agents fix names, types, units, dates, and time zones from every source before joining the tables on the right keys.

Fill the range.Keep it updated.

Agents check when new data appears and suggest a schedule. The first run fills your date range, then the dataset stays updated.

Download it.Query the rows.

Download the full dataset as Parquet, or query the latest rows as JSON through the same API.

Find the sources that fit the question.

Agents search existing datasets and public sources. They compare coverage, history, fields, format, and update frequency, then show which sources they would use.

DATA REQUESTHourly weather observations across the United States
SourceCoverageHistoryUpdate
NOAA ISDGlobal1901–nowHourly
IEM ASOSUS stations1928–nowMinutes
GHCNGlobal1763–nowDaily
MeteostatGlobal1901–nowDaily
ERA5Global grid1940–nowDaily

From messy sources to one live dataset.

Agents configure built-in readers for each API, file, and document. The pipeline downloads and parses the data, then runs again whenever the sources publish something new.

Connect directly

Open APIsLive endpoints
Keyed APIsAuthenticated
GraphQLQuery endpoints
WebSocketsLive streams
SQLDatabases
S3Object storage

Tables and records

CSVTables
JSONRecords
XLSXWorkbooks
XMLFeeds
ParquetColumnar files
NDJSONStreams

Documents and grids

ZIPArchives
PDFDocuments
HTMLWeb tables
GRIBWeather grids
NetCDFScientific arrays
GeoTIFFRaster maps

Turn mismatched tables into one consistent dataset.

Each column gets a clear name, type, and unit. Timestamps are converted to UTC. Every conversion is checked before the tables are joined. For temperatures, that includes both the scale and the offset.

SOURCE ROWS
temp_f67.1
local_time08:00 CET
stationPraha-Kbely
STANDARDIZE AND MATCH
67.1 °F19.5 °C
08:00 CET07:00 UTC
Praha-Kbelystation_id LKKB
READY DATASET
temperature_cfloat64 · °C
observed_attimestamp · UTC
station_idstring

Name the fields

Give each column a clear name, data type, and documented unit.

Standardize values

Convert measurements with the exact scale and offset required by their units.

Join the tables

Match rows on the right keys. A join is refused when declared units disagree.

Check the result

Verify row counts, date ranges, required values, and missing data.

Fetch the full history. Then keep it updated.

You choose the date range. Agents check when new data actually appears and suggest an update schedule. The first run fills the range, then the dataset stays updated automatically.

INITIAL FETCH LIVE UPDATES
START DATE FIRST BUILD

Use the same data in training and production.

Download the full history as Parquet to train models that predict real-world events, then query the latest rows through an API when those models need fresh data.

ONE LIVE DATASET europe_energy 18.4M rows · latest passing version
DOWNLOAD · FULL HISTORYeurope_energy.parquetParquet file
LIVE API · LATEST ROWSGET /datasets/europe-energy/rowsJSON response

Built for teams that work on data together.

Invite teammates and control who can build, manage, share, or only view datasets. Keep source credentials secure across the organization, and choose which datasets stay private, are shared by link, or become public.

Give everyone the right access.

Assign Owner, Admin, Editor, or Viewer roles. Changes and removals apply on the next request.

Private until you decide to share.

Keep datasets inside the workspace, share them by link, or publish them when you are ready.

Connect sources without passing keys around.

Credentials stay in Secret Manager and are released only to the exact ingest run that needs them.

Build the dataset your model needs.

Agents find the sources, clean and join the data, and keep it updated. Download the full history or query the latest rows through one API.

Talk to the team

Self-serve signup lives at app.mostlyright.md. This form is for everything it cannot do: procurement, invoicing, custom licensing, or anything you want to ask first. We reply by email.

I'm a

One email from the team. No spam.