How Data Journalism Investigations Start: A Practical Guide 2026

Data journalism investigations start with a question, not a spreadsheet. You pick one narrow thing a public body does, name what you would expect to see if it were working properly, and then go looking for records that would prove you wrong as readily as right. That order matters, because the reverse approach — download a big open data portal file and browse until something looks interesting — is how most beginners burn a month and end up with nothing publishable.

Most first investigations take six to twelve weeks from idea to a publishable draft, and the first two weeks are almost entirely writing and record requests, not analysis. A spreadsheet handles most of the early work; coding is a later unlock rather than a prerequisite.

Investigative data journalism is reporting in which a journalist finds, cleans and analyses a dataset to test a specific question, then publishes the findings with documented methods so readers can check both the finding and the limits of the evidence.

Table of Contents

What You Need

Before you request a single file, you want nine things sorted. Skipping any of them is how a promising idea dies quietly in week three.

  • A reporting question with a named actor. “Why is bus service unreliable” is a subject. “Which bus routes lose the most minutes per month, and who is accountable for them” is a question.
  • A preliminary hypothesis written in one sentence, including what would disprove it.
  • Candidate datasets and record types, including which agency holds each one.
  • Subject-matter expertise, even informal. Ten minutes with someone who works in the system will save you a week of bad measurements.
  • A spreadsheet tool. Excel or Google Sheets is enough for the first two thirds of most civic investigations.
  • A cleaning tool once the data is messy: OpenRefine handles inconsistent names and addresses better than anything else free.
  • Mapping and charting options — QGIS for geography, Datawrapper or Flourish for web charts.
  • Source contacts: the agency press office, the named custodian of records, and at least two people outside the institution.
  • Documentation habits: a source log, a file trail and a running note of every transformation you apply.
  • Legal and ethical checks around personal data, scraping terms and publication risk, decided before you have findings you feel attached to.

The last item is the one people defer. Once you have found something, you are no longer a neutral researcher, and the decision about whether an identifiable person appears in your chart gets much harder to make soberly.

Step-by-Step: How Data Journalism Investigations Start

Step-by-Step: How Data Journalism Investigations Start

The workflow below runs in eight stages, from naming the question to a publication plan. The order is deliberate: each stage produces something the next one needs, and skipping ahead usually means redoing work later.

1. Find a Focused Public-Interest Question

Start by writing down a broad subject, then narrow it until you can state who is affected and what you would expect to find. “Water quality” fails that test. “Which monitoring stations in one river basin recorded the most exceedance days last year, and is that downstream of a permitted discharge” survives it.

The fastest route to a good question is a dataset you already know exists, read backwards. Pick a recurring release — monthly crime figures, quarterly permit lists, annual spending returns — and ask what a sensible person would expect to see. Any gap between that expectation and the published pattern is a candidate.

Two other reliable sources: your own beat, where you already have the contacts and the context that makes a number meaningful, and other cities’ stories. A pattern found in a comparable jurisdiction is often a well-tested hypothesis rather than a guess.

Write the question down in the form you would say it out loud. If you cannot say it in one breath, it is still two questions wearing a coat.

2. Turn the Question Into Testable Hypotheses

A hypothesis is the part where beginners stall, because it sounds academic. It is simpler than that: it is your expectation plus the specific fields that would confirm or kill it.

Work backwards. List the variables you need — date, location, amount, category, demographics — then work out which body collects each one and under what name. “Delay” might live in a council’s committee minutes, “response time” in a call centre log, and “population” in a census release. Three variables, three sources, three different request formats.

Write at least two competing explanations. If your story is that agency A is systematically slower than agency B, the alternative is that they define “resolved” differently, or that case complexity differs by region. You cannot separate those without asking.

Also record what evidence would change your mind. A reporter who cannot say what would prove them wrong will keep finding reasons the data fits their hunch, usually without noticing.

3. Assess Data Quality, Access, and Ethics

Before you spend time on a dataset, find out whether it is fit for the question. Provenance first: who collects it, under what statutory duty, and how the collection method changed over the period you want to study. A definition change mid-series is the single most common way a good-looking trend turns out to be an artefact.

Then check coverage and completeness. Which places, dates or categories are missing? Missingness is rarely random — records tend to disappear where scrutiny is highest, which makes the gap itself reportable.

Access matters too. An agency that publishes a dashboard once a year and answers nothing is a different project from one with an open API and a named contact. Spend twenty minutes finding the custodian of records before committing, and send the request early while you still have other threads going.

On ethics: establish whether you will be publishing identifiable personal data, and at what threshold. Aggregation rules matter — a count of four or five in a small geography can identify someone by elimination even when no names appear.

4. Gather and Document the Evidence

Start a source log on day one and never stop updating it. One row per request: date sent, exact wording, who it went to, channel, deadline, response, follow-up dates. When a story takes six months, your memory of what you asked in week two is worthless.

Capture original files, not just the working version. Keep the raw download with its filename and access date, plus any metadata the portal exposes. Agencies revise datasets silently; an archived copy is the difference between a defensible story and a dispute you cannot settle.

Keep a transformation log too. Every join, filter, exclusion and reclassification gets a dated line: what you did, to which columns, and why. This is the raw material for the methodology note readers will hold you to.

Data rarely tells the whole story. Supplement it early with records — committee minutes, inspection reports, contract files — and with fieldwork. A pattern in a spreadsheet means nothing until you know what the underlying case looked like.

5. Explore the Data Without Locking In a Conclusion

Clean first: consistent date formats, consistent place names, duplicates removed, missing values flagged rather than silently zeroed. OpenRefine is the fastest route when names and addresses are the problem; a spreadsheet is fine for everything else in the early stages.

Then look for structure rather than a headline. Calculate rates, not raw counts — a city with more incidents may simply have more people. Establish a baseline and compare against it: median and range rather than average alone, because one enormous value drags an average away from anything typical.

Compare groups and geographies, and test whether your result survives different reasonable assumptions. If the pattern only appears when you exclude one month, or only when you use one of two definitions, you have found an artefact, not a finding.

Follow the results that contradict you. That is where the mistakes live, and where the second half of the story usually is.

6. Verify Anomalies and Seek Context

Every interesting outlier gets traced back to individual records until you know what it actually is. Is it a real case, a coding error, a duplicate entry, or a unit change from dollars to thousands that nobody announced? Most “smoking guns” in early analysis are one of the last three.

Benchmark against independent sources. If a dataset says inspector visit rates doubled, check against published inspection reports or workforce figures from another body. Two unrelated datasets agreeing is worth far more than one dataset examined thoroughly.

Then ask specialists what the data can and cannot establish. Academics, former officials and frontline staff will tell you, usually generously, which fields are trustworthy and which are aspirational. They will also tell you what a normal value looks like, which no dataset documents.

7. Develop Sources and the Human Story

Your analysis tells you who to talk to and what to ask. Someone named repeatedly in the outlier rows can explain the mechanism. Someone affected by the pattern can show what the numbers mean in practice. Someone who worked inside the agency can tell you which part of the process explains the gap.

Prepare properly before the interview. Ask about the mechanism, not the conclusion. Specific, evidence-based questions get specific, useful answers; leading questions get confirmations of what you already believe.

Data without a person attached is hard to care about, and a person without data is hard to generalise. The story works when both are doing work the other cannot do alone.

8. Build a Verification and Publication Plan

Build a Verification and Publication Plan

Have someone else re-run your key calculations from your raw files without talking to you. If they cannot reproduce your headline number, you have a documentation problem, and it will be worse in someone else’s hands.

Seek right of reply early, with specific questions and a real deadline, and record exactly what you asked and what you got. Institutions decline more often because the request was vague than because the story is protected.

Identify foreseeable harm in advance: reputational damage to identifiable people, retraumatisation of subjects, or re-identification through small counts. Decide what you will do about each before publication, not after complaints arrive.

Write the methodology note while the analysis is fresh: where the data came from, what you excluded and why, what the data cannot show. Publish the source data or as much of it as ethics allows, so others can check your work. Then read every sentence and check that each claim matches what your evidence actually supports.

Common Mistakes

Nearly every avoidable failure in a data investigation comes from rushing one of the first four steps. Here is what goes wrong and what to do instead.

Locking in the conclusion first. If you decide the answer before opening the file, every odd number becomes confirmation. Write the hypothesis down, including the disconfirming result, and put the file away for a day.

Choosing the wrong denominator. Comparing raw counts across populations of different sizes is the most common analytical error in civic reporting. Divide by the population you are claiming to have affected, and show the raw count alongside the rate so readers can judge the scale.

Cherry-picking the window. Choosing the start date that flatters your hypothesis is easy to do accidentally because dashboards default to the last twelve months. Fix your date range before looking at the numbers.

Truncated axes and mismatched charts. A bar chart with a y-axis starting at 800 turns a 2% wobble into a cliff. Bar charts must start at zero; for rates over time, use a line with a clear label on the axis break if you need one.

Causal overstatement. Data showing that two things move together does not show that one causes the other. Say what the evidence supports: that the pattern is consistent with a particular mechanism, then let your sources tell you whether the mechanism is real.

Under-checking sources. One anonymous file with no provenance is a rumour with a spreadsheet. Find at least one independent dataset that corroborates the core claim, and ask a specialist to attack your methodology rather than compliment it.

Inconsistent numbers between text and graphics. This happens when the text was edited after the chart. Reconcile both against the same source file, every time, and treat a mismatch as a blocking error.

Before you publish: have a second person reproduce your central calculation from raw files; read the methods note as if you were the subject’s lawyer; check that every chart has a source line, a unit and a date range; and confirm that your strongest claim is the one your evidence can carry, not the one that got the most reactions in a group chat.

Frequently Asked Questions

What is the best way to start a data journalism investigation?

Start with one narrow question about a specific public body, then work backwards: list the variables you would need, identify which organisation collects each one, and request those records before any analysis. Write your expectation and what would disprove it in a single sentence. This order keeps a hunch honest and tells you immediately whether the data you need actually exists.

Do I need programming skills to investigate data?

Not at the start. Spreadsheets handle the majority of early work in a civic investigation, including joins, pivot tables and basic rates. Learn OpenRefine for messy names and addresses, and SQL or Python later when you hit a genuine wall. Most beginners who stall do so because they skipped scoping and records work, not because they lacked code.

How do I know whether an anomaly in a dataset is meaningful?

Trace it back to the individual records first and rule out coding errors, duplicates and unit changes. Then check whether the same pattern appears in an independent source, and whether it survives reasonable alternative assumptions about dates and definitions. If it only appears under one specific choice you made, treat it as an artefact until someone else reproduces it.

What should I do if the data does not support my original hypothesis?

Treat it as information, not failure. Check first whether the definition, coverage or time period explains the gap, then look for the secondary finding you may have walked past. Write up what was ruled out and why. Publishing a well-evidenced question that turns out not to hold is a legitimate result, and the record of it protects your next pitch.

How should reporters handle confidential or sensitive source data?

Store it encrypted, keep access to the smallest possible group, and document provenance and acquisition dates. Do not download more than you need to test the hypothesis. Avoid publishing anything that could re-identify individuals through small counts or unusual attributes, and get legal advice before publication if exposure is foreseeable. Source protection matters more than scoops.

When is a data journalism investigation ready to publish?

When someone else can reproduce your central number from your raw files without asking you a question, you have sought right of reply with specific questions, and every sentence in the draft matches what the evidence supports. Publish with a methods note explaining what the data cannot show, plus source data where ethics allow. Urgency is not one of the criteria.

Conclusion

Investigations that stall almost always stalled at the same point: nobody narrowed the subject into a question with a named actor and a disconfirming test. Fix that and the rest is ordinary, tedious, learnable work. Pick one public body, one focused question, one small recurring dataset and write the hypothesis on paper. Send the records request this week, and let the analysis follow the data rather than lead it.

Leave a Comment