To measure impact of a civic app, you decide which public outcomes it is supposed to change, capture a baseline before launch, then track a small set of indicators that would prove or disprove that change. Downloads and monthly active users are activity counts, not impact. The real test is whether a service request got fixed faster, whether people who had never contacted the city before did, and whether anything in the budget or a service standard moved as a result.
Most measurement guides for government programmes are written for social programmes, where outcomes are vaguer. A civic app is a closed loop: report, route, resolve, trust. That loop gives you unusually clean measurement points, and the framework below is built around them.
Set aside two or three working days for the first pass. The ongoing reporting rhythm after that is a few hours a month if the instrumentation is set up properly.
Table of Contents
- What You Need
- How to Measure Impact of a Civic App: Step-by-Step
- 1. Define the App’s Public Outcome
- 2. Build a Theory of Change
- 3. Establish the Baseline and Targets
- 4. Choose Impact and Guardrail Metrics
- 5. Collect Reliable Evidence
- 6. Analyze Change and Segment Results
- 7. Decide What to Scale, Improve, or Stop
- Measurement Cadence for a Team With No Analyst
- Common Mistakes
- Frequently Asked Questions
- What is the difference between output and outcome metrics for a civic app?
- How do I set a baseline before a civic app has launched?
- How many survey responses do I need for a city app claim?
- What are vanity metrics in civic technology?
- How do you measure whether a civic app actually made things better?
- How often should a city report app impact metrics?
- Conclusion: What to Do First
What You Need
You need a named owner for the measurement, and it does not have to be an analyst. A digital services lead who can pull a number and write two paragraphs is enough to start.
Then you need the following.
- One public outcome, written down. A sentence describing the change in the city, not in the app. “Pothole reports get repaired” works. “More people report issues” does not, because that is a consequence you can control rather than a change you caused.
- Baseline evidence. Twelve months of pre-launch history if you can get it: report volumes, median resolution time, how many arrived by phone, what share of residents had contacted the city in the past year. Without this, a rise in activity after launch tells you nothing.
- Data sources you can actually access. The app’s own event data, the work-order system, the call centre log, the council CRM. Write down which owner will approve each pull and how long that takes.
- Measurement tooling. The analytics you already have is usually enough for the first cycle. A separate warehouse, a survey platform and a shared dashboard are useful, not mandatory.
- Consent and privacy materials. A plain-language notice covering what you collect, why, how long you keep it, and how someone opts out. Public-sector analytics teams usually already have a template; ask before building a new one.
- A reporting period and a named audience. A quarterly report to a service committee and an annual report to council are different documents with different lengths. Decide which one you are writing before you choose the metrics.
One more thing worth writing down at the start: what you would do differently depending on the result. If the answer is “nothing, we would keep the app regardless”, you have built a monitoring exercise, not an evaluation.
How to Measure Impact of a Civic App: Step-by-Step

1. Define the App’s Public Outcome
Start with one sentence a councillor would recognise as a city problem, not a product problem. “Fewer streets go unrepaired for more than a month.” “Fewer residents miss a bin collection because they do not know the schedule.” “More people find out what a council decision is before the vote, not after.”
Then list the measurable changes underneath it. Three to five is usually the working limit. For a pothole-reporting app, that could be median days from report to repair, the share of reports closed within the published service standard, the share of reports that come with a usable photo so a crew finds the right location first time, and the share of reports coming from neighbourhoods that were previously quiet.
Keep adoption out of the outcome list. How many people downloaded the app is something you decide, so it cannot prove the app mattered.
2. Build a Theory of Change
A theory of change is a chain you can draw on one page: inputs, activities, outputs, outcomes, impact. The app is one input. The activities are what the city does with it. The outputs are requests created, routed, repaired, and confirmed. The outcomes are faster repairs and fewer repeat calls. The impact is a streets budget that spent more of its time on repair and less on chasing information.
The useful part is the last third of that page, where you write what the app cannot reach. A reporting app cannot fix a shortage of paving crews, and it cannot change whether a resident trusts the council. Saying so up front protects you later, when someone asks why a strong month did not reduce the backlog.
If the chain has more than five links between a user action and a public outcome, you are probably claiming more than the app can carry. Cut it down until each link has an owner.
3. Establish the Baseline and Targets
Capture the baseline before you write a single target, and write the date on it. Twelve months is better than three because seasons distort road and waste reporting badly. If the app has already been live, reconstruct the pre-launch period from archived data and label it as reconstructed.
Record the numbers as they are, not rounded. Median days to resolution, not average, because a handful of very old cases drag an average into uselessness. Percentages, not counts, so a change in total volume does not read as a change in performance.
Then set thresholds rather than targets. A threshold states what evidence would trigger action: “if median resolution improves by 15% or more and equity of participation does not fall, we scale.” A target invites you to move the goalposts when the data is inconvenient. A threshold is harder to argue with, and it makes the eventual decision look reasoned rather than political.
Where you can, note a comparison group. Two similar districts, one with the app and one without, is a rough but honest control. Failing that, use the pre-launch period and say plainly that season, weather and staffing all moved during it.
4. Choose Impact and Guardrail Metrics
Impact metrics show whether the outcome moved. Guardrail metrics show whether you got there by damaging something else. Most civic app dashboards are all guardrail and no impact, which is why they end up as download charts.
The metric tree below organises what to track into five families. The last two are the ones most city dashboards skip.
| Family | Example metric | How to collect it | What it proves |
|---|---|---|---|
| Adoption | Share of residents using the app at least once in a year | Unique accounts, compared against the service-area population | Whether the app reached beyond the existing phone-and-email users |
| Engagement | Reports per active user, 90-day retention, completion rate | App event log | Whether people come back or abandon the flow halfway |
| Resolution | Median days to resolve, share closed within the service standard, reopen rate, routing accuracy | Work-order and CRM records joined to the request ID | Whether the closed loop actually closes |
| Governance | Share of reporting that changes a work order, contract, budget line or published standard | Manual tracking against a named decision-maker | Whether the app’s output reaches a decision, not just a queue |
| Outcome and equity | Repeat complaint rate, satisfaction with getting heard, participation gap by neighbourhood income and age | Repeat-call analysis, survey, open data by area | Whether the public result improved, and for whom |
Two metrics deserve special attention because they are what separates a civic app from a commercial one. Routing accuracy, the share of reports that reach the right team first time, is a cost metric hiding inside a quality metric: misrouted work burns a crew’s day. And participation equity, measured by neighbourhood income and age, is the one that decides whether a service improvement widened or narrowed a gap.
The POPVOX CivX Metrics Toolkit argues that customer satisfaction is the wrong frame for government services. Its families are worth borrowing: civic efficacy, meaning whether people believe their voice and their actions matter, civic wayfinding, meaning whether they can find the right service and the right department, and trust in the institution. Those three map cleanly onto app instrumentation later in this process.
5. Collect Reliable Evidence
No single source carries the argument. Use four, and know what each is weak at.
| Method | What it is good for | What it cannot do | Typical lag |
|---|---|---|---|
| App analytics | Reach, engagement, funnel behaviour, cost per completed request | Anything about people who never installed the app | Immediate |
| Operational records | Resolution, service-standard compliance, reopen rate, unit cost | Whether the resident is satisfied with the outcome | Weeks |
| Survey | Trust, feeling heard, satisfaction, demographic breakdowns | Explaining why a number moved; low response from the least engaged users | One cycle |
| Interviews and listening sessions | Why a flow confused people, why trust fell, what happens after a report | Rates, trends, or anything you would put a number on | One cycle |
Document the source, the collection method, the response rate and the known gaps for every figure you plan to publish. This is dull work and it is the difference between a measurement report and an anecdote. A 40-response in-app survey can tell you what went wrong in a flow; it cannot support a claim in front of a council committee, and saying so in the footnote is better than being asked in the room.
Recruit beyond your user base. People who gave up are the ones with the useful information, and they are usually not answering your in-app prompt.
6. Analyze Change and Segment Results
Compare the result to the baseline first, then ask whether the change is big enough to be more than weather. A 9% improvement in median resolution time in a year when crews were reallocated is not evidence. A 40% drop in repeat calls after a routing change, sustained across two quarters, is at least worth investigating.
Segment before you conclude. Break the results out by neighbourhood, by age band, by language and by channel. Two patterns hide in every average: the app working well only in the areas that already complained loudest, and the app quietly replacing a phone line for the people who are comfortable online while the rest wait longer. Neither shows up in a single headline number.
Then be careful with cause and effect. Most of the time you can claim contribution, not attribution. A pothole repaired after a report is a plausible link, not a proven one, since some of those streets were already on a repair schedule. Write “the app contributed to” rather than “caused”, unless you ran something close to a proper test: a randomised rollout district by district, or a staggered start where half the city got the app a quarter before the other half. Both are achievable, and both are almost never done.
7. Decide What to Scale, Improve, or Stop
Now apply the thresholds you set in step three, before you look at the results again. If median resolution improved past the threshold and participation equity held flat, the decision is scale, and the next line of the report is which features to expand, not how well things went.
Name what caused the change where you can. If one report category, one channel or one routing rule accounts for most of the improvement, you have something to expand and something to retire. An app whose gains come from a single feature is one budget cycle away from losing it.
Record the trade-offs honestly. A fast closure rate achieved by auto-closing requests after seven days will look excellent in every dashboard and will be reversed by the first ward councillor who sees a fake statistic. Write down what the app did not fix, what it made harder, and who owns each next action with a date against it.
Measurement Cadence for a Team With No Analyst
Not every metric needs the same clock. Weekly review of a handful of operational numbers catches a broken workflow in days. Quarterly is when you bring in resolution, cost and the survey. Once a year is when you compare against baseline, re-run the equity breakdown and decide whether the outcome is still the one you are aiming at.
Keep the weekly list under eight numbers. A dashboard nobody opens is the most common failure in municipal measurement, and it usually starts with too much, not too little.
Common Mistakes
Treating downloads as impact. An app can be downloaded forty thousand times and change nothing, and the download count says nothing either way. MIT GovLab’s practitioner guide, “Don’t Build It”, names this pattern directly: metrics that make the technology look good without reflecting real impact are vanity metrics, and reporting them is worse than reporting nothing, because they occupy the space where evidence should be. Fix it by pairing every activity number with the outcome number it is supposed to influence.
Reporting outputs and calling them outcomes. “4,200 reports submitted” is an output. “Median repair time fell from 31 to 19 days” is an outcome. The first tells you the app is used, the second tells you the city is better. If your report has no outcome number in it, you have an activity log.
Choosing metrics after seeing the data. A metric set picked in month eight is always picked to flatter month eight. Fix it by writing the set down in step four and, if you must add one later, recording why in the same sentence.
Skipping the comparison group. Every city metric trends upward in good years. Without a pre-launch baseline or a not-yet-covered district, a rising resolution rate is indistinguishable from rising expectations. Fix it by starting a rollout in a defined area rather than launching citywide on day one.
Hiding the bad result. A report where every number improved gets disbelieved the first time it does not, and the team loses the credibility it needs for the next funding round. Publish the metric that went the wrong way, with the same prominence.
Collecting feedback without consent or representation. Asking residents to rate their government experience inside a service flow, with no notice and no opt-out, damages the trust you were trying to measure. Fix it with a plain notice, an easy decline, and a small incentive for people who take part in research rather than a service interaction.
Over-reading small samples. Forty responses describe forty people. They are a good warning system and a bad claim. Say which one you are using.
Frequently Asked Questions
What is the difference between output and outcome metrics for a civic app?
An output counts what the system did: reports submitted, crews dispatched, notifications sent. An outcome counts what changed because of it: median repair time, fewer repeat calls, a ward where residents report more problems because they trust it will be fixed. Outputs are easy to measure and useful for operations, but they only prove the app is being used. Outcomes are harder, slower and require data from outside the app, which is why a report without them reads as an activity log.
How do I set a baseline before a civic app has launched?
Pull twelve months of history from the channels the app will replace: call centre logs, email, counter intake and the work-order system. Record report volume, median time to resolve, the share closed within the existing service standard, and how many residents contacted the city at all. Where the app is already live, reconstruct the pre-launch period from archived records and label it as reconstructed. Twelve months is the minimum that lets you separate a real trend from the season.
How many survey responses do I need for a city app claim?
Enough that you would be willing to defend the number in public, which is a much higher bar than the formula suggests. Roughly 400 responses gives a margin of error near plus or minus five percentage points at 95% confidence, and you need more for any subgroup breakdown by neighbourhood or age. Below that, report what you heard as themes from interviews and say plainly that the survey is directional. An in-app prompt sent to happy users will never produce a representative sample on its own.
What are vanity metrics in civic technology?
They are numbers that make the project look successful without measuring whether anything changed for the public. Downloads, registrations, social mentions and session counts are the classic examples, and MIT GovLab’s practitioner guide Don’t Build It warns against them directly. A vanity metric is worse than no metric, because it takes the slot where evidence should go and it survives scrutiny easily. Pair every activity number with the service or outcome number it is meant to influence.
How do you measure whether a civic app actually made things better?
Compare against a baseline captured before launch, use a comparison area where you can, and be honest that you are describing contribution rather than proof. The strongest design is a staggered or district-by-district rollout, so some parts of the city experience the app months later than others. Without that, describe results as contribution, note what else changed at the same time, such as budgets, staffing or policy, and avoid causal language in anything a decision-maker will quote.
How often should a city report app impact metrics?
Weekly for a short operational list of under eight numbers, so a broken workflow shows up in days rather than a quarter. Quarterly for resolution, cost and a short survey, which is what service committees should see. Annually for the baseline comparison, the equity breakdown by neighbourhood and age, and the decision about whether the original outcome is still the right one. Teams that report everything weekly end up reporting nothing by month three.
Conclusion: What to Do First
Before you build a dashboard or commission a study, write three things on one page: the single public outcome you are trying to change, the baseline number for that outcome as it stands today, and three indicators that would prove or disprove progress. Everything else, the event taxonomy, the survey, the equity breakdown, the quarterly report, comes after those three lines exist.
If you do that much, you will already be ahead of most city apps in 2026, and you will know within a quarter whether the app deserves to exist.


