Fetch answers from
messy data
csv.dog, pdf.dog, and xml.dog turn the files nobody wants to open into datasets you can question, query, and pipe. Usage-based pricing — you pay when it's useful, and nothing when it's not.
The Problem
The answer is in the file.
Getting it out is the job.
Every team has a folder of CSVs, PDFs, and XML exports that contain exactly what they need to know — wrapped in formats that fight back. So the question waits on a cleaning script, a data engineer, or a subscription nobody wants to add.
The same cleaning script, rewritten forever
Encodings drift, dates arrive in twelve flavors, numbers show up wearing currency symbols. Every analysis starts by rediscovering the same problems and writing the same throwaway script nobody can reproduce later.
Files too big for spreadsheets
Past a few hundred thousand rows, spreadsheets stop being tools and start being crash reports. The data isn't too big to matter — just too big to open.
Formats built for printing, not questions
PDFs and XML hold real data hostage behind layout and nesting. You can see the numbers right there on the page. You just can't query them.
Tools that bill you for not using them
Subscription data platforms charge the same on the month you use them daily and the month you forget they exist. The price has nothing to do with the value you got.
Products
Every format gets
its own dog
One tool per format, all with the same promise: upload your data, ask a question, fetch the answer — and pay only for what you use.
csv.dog
Upload one CSV or a hundred, ask questions in plain language, fetch answers as shareable tables and charts. Build pipelines that clean, join, and score the next upload automatically.
pdf.dog
The data is in there — trapped behind the layout. Pull tables and fields out of PDFs into datasets and pipelines you can actually query.
xml.dog
Deeply nested, rigorously specified, still a pain. Flatten XML feeds and exports into queryable data and repeatable pipelines.
The Solution
Point a dog at it
Each tool in the family does one thing: take a format developers dread and make it answer questions. Upload the file, ask in plain language, and turn what worked into a pipeline you can run on the next one.
Bring the messy file
One file or many, combined into a single queryable dataset. No schema ceremony, no import wizard marathon — the file becomes data you can work with.
Question it in plain language
Ask the way you'd ask a teammate. Get tables, numbers, and visualizations back — with real SQL available whenever you want to be precise.
Get answers you can share
Results come back as clear, shareable tables and charts you can hand to a teammate — not screenshots of a terminal session.
Make it repeatable
Teach it once. Reusable pipelines clean fields, derive columns, score rows, and join other data — so the next upload arrives already prepared.
How It Works
Upload. Ask. Fetch.
The whole product in three steps — and a fourth that makes the next file effortless.
Upload your data
Drag in a file — or a pile of them. Multiple files combine into one queryable dataset, so you can ask across all of them at once.
Ask a question
Plain language first, SQL when you want it. No modeling step, no BI onboarding, no waiting for someone with database credentials.
Fetch the answer
Answers come back as clear tables and visualizations you can share with the people who asked — ready to use, not ready to reformat.
Build the pipeline
Turn the cleanup and derivations that worked into a reusable pipeline. The next upload gets cleaned, enriched, and joined before you ask anything.
Capabilities
Built for messy data,
not demo data
Real-world files are the home turf — drifting schemas, sentinel nulls, ambiguous dates, and all. Every capability assumes the data arrives imperfect.
Plain-language questions
Ask the way you'd ask a teammate — no query builder, no modeling step. The answer comes back as a table or chart, not a syntax error.
Real SQL when you want it
Plain language is the front door, not a ceiling. Drop into SQL whenever precision matters, on the same dataset.
No row is ever dropped
Values that fail parsing aren't silently nulled or skipped. They're quarantined with the raw bytes and the reason, so you can fix them and replay.
Reusable pipelines
Clean fields, derive columns, score rows, join other data — then save it. Every future upload runs through the same steps before you ask a thing.
Many files, one dataset
Combine a folder of exports into a single queryable dataset and ask across all of them — no manual concatenation, no vlookup gymnastics.
Shareable tables & charts
Answers become clean tables and visualizations you can share with the person who asked — built for handing off, not for exporting and reformatting.
Pricing
No subscriptions.
Usage, and nothing else.
You pay for the bytes you store and the bytes your questions search. That's the entire model. If a tool is useful, we see usage — if it isn't, we don't, and you owe us nothing for the quiet months.
- A flat monthly fee, used or not
- Per-seat licenses that punish sharing
- Feature tiers designed around upgrades
- “Contact sales” standing between you and a price
- Cancellation flows built to be forgotten
- $0.05 per queryable GB per month
- $0.005 per GB searched by your queries
- A dataset nobody queries costs almost nothing
- Free local desktop app for data that stays home
- No seats, no tiers, no contracts
The model is deliberately boring: billing follows the bytes you store and search. Our incentive is to make each tool worth coming back to — not to make you forget to cancel.
Who It's For
For people with a file
and a question
You don't need a data team, and you shouldn't have to become one to get an answer out of a file.
Analysts and operators with a file in hand
You have the file and the question — what's missing is the hours of cleanup between them. Upload, ask, and get on with the decision instead of babysitting a spreadsheet.
Developers building on messy sources
Public filings, partner feeds, billing extracts — recurring inputs that deserve better than a cron job full of regret. Turn the cleanup into a reproducible pipeline that runs on every new file.
Teams whose data can't leave the building
The free desktop app runs the same pipeline locally, so sensitive files get queried on your machine — not uploaded to prove a point. Cloud is there when you want sharing and scale.
FAQ
Frequently asked questions
Early Access
Bring the file.
We'll fetch the answer.
csv.dog is in development with early users now — pdf.dog and xml.dog are next. If you have a folder of files and a backlog of questions, we'd like to hear about both.
No subscription to start, no subscription ever. You pay for usage — that's the whole deal.