Real-time transit data feeds work as a single chain: an onboard GPS unit pings the agency’s server, a dispatch system joins those pings to the published schedule, a feed builder serialises the result as a GTFS-Realtime protobuf message, and apps and maps poll it to show riders where the bus is and when it will arrive. Everything below unpacks that chain, piece by piece.
I have built and debugged city apps that sit on top of these feeds, and the thing that surprises new developers most is how little of the “real time” part happens on the vehicle itself. The bus reports where it is. Everything else — what that means, when the rider’s stop is coming, whether the trip is cancelled — is added by systems on the ground.
Written for 2026, this guide walks the whole path: what counts as real-time transit data, where it comes from, how predictions are made, which formats carry it, and what breaks in practice.
Table of Contents
- What Are Real Time Transit Data Feeds?
- What exactly is real-time data in this context?
- GTFS and GTFS-Realtime are two different things
- How Real Time Transit Data Feeds Work From Source to Screen
- Where Does the Transit Data Come From?
- Automatic vehicle location
- Automatic passenger counters and fare systems
- Traffic signals and priority systems
- Operator applications
- Third-party and derived sources
- What Happens Inside the Data Pipeline?
- How Do Transit Arrival Predictions Get Calculated?
- Schedule-based estimation
- Position and headway learning
- What goes into a prediction
- Why two apps show different times for the same bus
- Which Data Formats and APIs Do These Feeds Use?
- How Do Transit Apps Receive and Display the Data?
- Polling and conditional requests
- Caching and joining
- Map matching and display
- Arrival presentation
- Graceful degradation
- How Is Feed Quality Measured and Maintained?
- What Are the Main Reliability and Privacy Challenges?
- When real-time data goes wrong
- Serving, licensing and rate limits
- Location privacy
- Frequently Asked Questions
- What is GTFS-Realtime?
- How do transit apps know where the bus is?
- How often is a GTFS-Realtime feed updated?
- Why do real-time transit apps show wrong times?
- Do I need an API key to access a GTFS-Realtime feed?
- What is the difference between TripUpdate and VehiclePosition?
- Conclusion
What Are Real Time Transit Data Feeds?
A real-time transit data feed is a live, machine-readable stream of where transit vehicles are, when they are expected to arrive, and what is disrupting service right now. Agencies publish them so apps, maps, stop displays and call centres show current conditions instead of a printed schedule.
The useful distinction is between what was observed and what was planned. A vehicle position ping is an observation: it happened, and it has a timestamp. A predicted arrival is a calculation. A scheduled arrival is a plan from a file published weeks in advance. Mixing them up is the source of most rider confusion.
What exactly is real-time data in this context?
Data reflecting the current state of a transit system, updated at intervals of seconds to minutes, as opposed to a static published schedule that changes only when the agency republishes it. A GPS coordinate, a predicted arrival minute and a detour notice all qualify, because all three change within minutes of the event they describe.
GTFS and GTFS-Realtime are two different things
Most agencies publish both. GTFS is the static schedule — a ZIP of plain text files describing routes, stops, trips and timetables. GTFS-Realtime is the live layer, published at a separate permanent URL. Consumers almost always load both and join them at runtime.
| Aspect | GTFS (static) | GTFS-Realtime |
|---|---|---|
| What it holds | Routes, stops, trips, timetables, fares | Live positions, arrival predictions, service alerts |
| Format | Zipped CSV text files | Protocol Buffers (binary), sometimes JSON variants |
| Delivery | One download, refreshed occasionally | Continuously updated endpoint, polled or subscribed |
| Typical change rate | Daily, weekly, or per service change | Every few seconds |
| Stable identifiers | route_id, trip_id, stop_id | References the same IDs plus live entity IDs |
That last row matters more than it looks. A GTFS-Realtime TripUpdate does not describe a route from scratch. It points at a trip_id that already exists in the static feed, and fills in the parts that have changed.
How Real Time Transit Data Feeds Work From Source to Screen

Six steps, end to end. This is the version worth memorising, because every failure in a real-time transit app can be traced back to one of them.
- Collect position pings on the vehicle through automatic vehicle location hardware.
- Receive them at the agency server, where pings from every vehicle land in one stream.
- Match each ping to a trip in the static schedule using the dispatch system.
- Estimate future arrivals at each remaining stop from current position and history.
- Serialise and publish TripUpdates, VehiclePositions and Alerts as one protobuf message at a public URL.
- Poll, cache and render in a consumer backend, then show the result on a rider’s phone or stop display.
Step five is where a lot of agencies buy software rather than build it. Commercial vendors such as Cubic and the NextBus platform, and open stacks such as OneBusAway, sit between the agency systems and the published feed. Nothing about the format changes; the vendor just handles the serialisation and hosting.
Step six is where the rider actually meets the data, and where the remaining sections of this guide pick up: how a feed is built, how predictions are made, and what happens when something in the chain goes quiet.
Where Does the Transit Data Come From?
Five sources feed most transit agencies, and each one has a different failure mode. Knowing which is which tells you why a feed degrades in the way it does.
Automatic vehicle location
AVL is the backbone. A GPS or GLONASS receiver on the roof, a cellular or radio link back to the agency, and an antenna. The unit reports position, heading, speed and odometer on a fixed cadence, usually every few seconds to every 30 seconds. Where a vehicle runs in a downtown canyon or a covered terminal, the pings degrade or stop entirely, which is why good systems fall back to dead reckoning using speed and heading.
Automatic passenger counters and fare systems
APC sensors at the doors and smart fareboxes with their own telemetry both emit boarding and alighting counts. That is where occupancy and crowding figures come from, and it is the only source that tells you how full the bus is rather than where it is. Accuracy is good but not perfect; door sensors misread when a wheelchair ramp deploys or a passenger boards twice.
Traffic signals and priority systems
Transit signal priority talks to intersections, and where an agency exposes that data downstream it explains a jump in travel time. Rarely published, but useful internally for reconstructing why a run got slower.
Operator applications
Driver tablets and run-control software generate events that GPS cannot infer: a manual detour, a trip dropped, a bus short-turned, an unscheduled trip added at the end of the line. These events are what turn a position stream into a service picture.
Third-party and derived sources
Some feeds are repackaged rather than produced by the agency, and app backends that scrape an agency website instead of parsing a feed are common enough that two apps can disagree about the same route. Bikeshare and micromobility use a completely different standard, GBFS, which covers station and vehicle availability rather than schedules.
What Happens Inside the Data Pipeline?
Inside the agency, between raw pings and a published message, seven distinct jobs happen. Skipping any of them is how feeds end up with phantom buses.
- Ingestion. Pings arrive over cellular or a vendor network and land in a queue. Many agencies keep a short buffer so a brief network outage does not create a gap in the stream.
- Validation. Coordinates outside the service area, impossible speeds, impossible timestamps and malformed packets are dropped or quarantined. Cheap checks catch most GPS noise.
- Deduplication. Retransmissions are common. Identical consecutive pings are collapsed so a stationary bus is not treated as teleporting.
- Geospatial processing. Raw coordinates are snapped to the road network, matched to the shape of the route, and checked for plausible heading. This map-matching step is widely regarded by developers as the hardest part of the whole pipeline.
- Trip matching. The snapped position plus time is joined to the static schedule to work out which trip the vehicle is running and how far along it is. Frequency-based trips, which have no fixed timetable, complicate this because any trip instance may match.
- Prediction. Remaining stop times are estimated and attached as stop-time updates, along with delay, cancellation and detour flags.
- Publication. The message is serialised and written to a URL that consumers poll. Alerts are merged in from the agency’s alert system at the same moment.
Trip matching is where a lot of confusion lives for newcomers. A bus that starts its run late may be matched to the wrong trip_id entirely, and everything downstream inherits the mistake. Consumers who join GTFS-Realtime to the static feed by trip_id alone will show a vehicle on the wrong route with confident-looking times.
How Do Transit Arrival Predictions Get Calculated?
Predictions come in two broad flavours, and most agencies use a mix of both.
Schedule-based estimation
The simplest approach reads the vehicle’s position against the timetable. The system knows how far the bus has travelled along the route and how long the remaining stops usually take, then subtracts. It is cheap, it works on nearly every route, and it is wrong in exactly the predictable ways: rush hour, a stalled traffic incident, a long dwell at a busy stop.
Position and headway learning
The better approach uses where the vehicle has actually been. On routes where service runs at a headway rather than a timetable, there is no scheduled arrival to fall back on, so the system predicts from recent spacing between vehicles. If the bus ahead is slow, that wait carries backward. On scheduled routes, agencies blend this with historical travel-time patterns learned per time-of-day and per segment.
What goes into a prediction
- Current position along the route shape and the last known heading.
- Recent travel speeds, which absorb traffic faster than a static profile does.
- Dwell time at the upcoming stops, estimated from boardings at previous stops.
- Service pattern, including frequency-based trips with no timetable at all.
- Schedule adherence, when the trip is running on or off its published time.
Why two apps show different times for the same bus
This comes up constantly on developer forums and rider forums alike, and it has a short answer: they are not reading the same snapshot. One app may cache for a minute, another for ten. One may be connected to the agency’s feed while another uses a downstream vendor’s copy. If the underlying feed is accurate, the disagreement is the app’s cache, not the bus.
For riders, the practical habit is to compare two apps at the moment you would decide whether to walk. When they disagree by more than a few minutes, one of them is not polling the feed at all.
Which Data Formats and APIs Do These Feeds Use?
GTFS-Realtime dominates, but it is not the only option, and the delivery method matters as much as the format.
| Format | Delivery method | Strength | Typical use |
|---|---|---|---|
| GTFS-Realtime | HTTP polling of a protobuf endpoint | Single global standard, strong tooling | Most North American and increasingly global agencies |
| GTFS-Realtime JSON | HTTP polling | Readable without a protobuf compiler | Debugging, small tooling, quick prototypes |
| REST APIs | Request and response per query | Easy to call from anything | Agency-specific or vendor endpoints, stop arrival lookup |
| Webhooks and subscriptions | Push to a registered consumer | Lower latency, less wasted polling | Large consumers and operations dashboards |
| SIRI | XML or JSON over SOAP or REST | Strong European standard, multimodal | UK, France and parts of the EU rail and bus |
| TAP TSI | XML-based operational data exchange | Deep operational detail, EU standard | European rail operators and infrastructure managers |
| GBFS | HTTP polling of JSON | Simple, well-maintained, bikeshare-specific | Shared bikes and scooters, not scheduled transit |
Under the hood, GTFS-Realtime is Protocol Buffers, Google’s binary serialisation format. A message contains a FeedHeader with a timestamp and version, then a list of FeedEntity objects, each holding a TripUpdate, a VehiclePosition or an Alert with a unique id. Optional fields in protobuf are genuinely absent rather than set to zero, so code that reads a field without checking whether it is present will happily process an empty value as if it were data.
For developers, the practical starting point is the official language bindings, published as gtfs-realtime-bindings. Validators from the same project will tell you whether a feed you are about to depend on is structurally sound, which is a cheap test before you build anything.
How Do Transit Apps Receive and Display the Data?

Consumption is a small engineering problem that decides whether your app feels trustworthy.
Polling and conditional requests
Consumer backends fetch the feed on a timer, typically every 15 to 60 seconds, often staggered so ten thousand clients do not hit the agency at the same instant. Sending If-Modified-Since or checking the ETag lets the agency answer “nothing changed” cheaply, which matters because many small agencies pay per gigabyte of egress.
Caching and joining
Raw feed data is cached for a short window, then joined against the downloaded static GTFS to resolve trip names, stop names and route colours. The cache is what lets an app answer instantly while the upstream feed is unchanged.
Map matching and display
Vehicle coordinates are snapped to the route geometry from shapes.txt, then drawn with the route colour and the agency’s own iconography. Where no shapes exist, a consumer has to fall back on road matching, which is noticeably worse at intersections.
Arrival presentation
The arrival list is the part riders judge you on. Showing scheduled and predicted times side by side, marking a prediction as older than a couple of minutes, and naming the vehicle or the trip so two buses on the same route can be told apart all do more for trust than any amount of map polish.
Graceful degradation
When updates stop, say so. A greyed-out last-known position with a visible last-updated time is far less damaging than a stale prediction that keeps counting down as if it were live. Riders forgive estimates; they do not forgive a frozen screen pretending otherwise.
How Is Feed Quality Measured and Maintained?
Feed quality has a handful of measurable dimensions, and agencies that publish them earn more credibility than agencies that claim everything is fine.
- Freshness. How old the newest record is. GTFS-Realtime best practice recommends refreshing at least every 30 seconds, keeping trip updates and vehicle positions no older than 90 seconds, and no older than 10 minutes for alerts.
- Completeness. Whether every active trip has a vehicle position, and whether positions exist during the hours of service. A feed that is perfect at rush hour and empty at 10pm is technically compliant and practically useless.
- Accuracy. How far predicted times drift from actual arrival once measured after the fact. A few minutes early in traffic is realistic; systematic errors of fifteen minutes are a data problem.
- Latency. The gap between the vehicle reporting and the moment the update is visible in an app.
- Coverage. Which routes and modes actually publish positions, as opposed to trip updates only.
Monitoring happens at two levels. The agency watches its own pipeline and alerts on a stalled feed or an unexplained drop in vehicle count. Consumers watch too, and should: a synthetic check that fetches the feed and counts fresh entities is a few lines of code, and it is the only way to learn that a feed quietly went stale at 6am on a Sunday.
Riders report bad data too. Agencies that publish a way to flag an incorrect position, and that visibly act on it, get treated more fairly when the inevitable error happens.
What Are the Main Reliability and Privacy Challenges?
When real-time data goes wrong
- Stale feeds. The most common failure by far. Positions stop updating while the app keeps rendering the last known state, so a bus appears frozen in place.
- Phantom and missing vehicles. A unit keeps reporting after its run ends, or disappears mid-route because of a cellular dead zone. Riders describe this as the bus vanishing.
- Duplicated vehicles. Two pings matched to the same trip, often after a garage or yard antenna picks up several buses at once.
- Orphaned identifiers. Trip ids in the live feed that no longer exist in the static feed, which happens after a schedule change has not propagated to the app.
- Missing alerts. Service is disrupted and the trip updates go stale, but no alert was ever published. Riders see a wrong time and no explanation.
- Schedule changes. The static feed updates faster than consumers re-download it, so live updates reference stops that no longer exist in the app’s copy.
The practical defence against all six is the same: validate freshness on every read, degrade visibly, and never show a prediction without the timestamp that produced it.
Serving, licensing and rate limits
Public agency feeds are usually free and unauthenticated, which is the norm. The moment you add a vendor gateway, API keys appear, quotas appear, and a key embedded in a mobile app becomes a shared secret. Budget for a backend proxy rather than calling keyed endpoints from the device. Some agencies also restrict redistribution or commercial use of their data, so read the terms before you ship.
Location privacy
Precise vehicle positions are location data about passengers as much as about buses. A bus dwelling near a clinic, a school or a house at 11pm is telling you something about someone. Publishing only vehicle positions and never passenger-level traces is the normal line, and keeping raw pings in short-lived caches rather than long-term stores limits the exposure. If your product builds history from live positions, that is a deliberate decision worth documenting.
Frequently Asked Questions
What is GTFS-Realtime?
GTFS-Realtime is a specification for publishing live transit information alongside a static GTFS schedule. A feed is a Protocol Buffers message, served over HTTPS at a permanent public URL, containing three entity types: TripUpdate for arrival and departure predictions, VehiclePosition for live location, and Alert for disruptions. It was developed in the United States and is now used by agencies worldwide.
How do transit apps know where the bus is?
An onboard GPS or AVL unit reports position, heading and speed to the agency server on a short cycle. The dispatch system matches each ping to a scheduled trip, and a feed builder publishes it as a VehiclePosition in the GTFS-Realtime feed. Apps and maps poll that feed every 15 to 60 seconds and draw the vehicle on a map. Nothing is tracked inside the app itself unless the vendor operates its own hardware.
How often is a GTFS-Realtime feed updated?
Best practice, as set out in the GTFS-Realtime documentation, is to refresh the feed at least every 30 seconds and to keep trip updates and vehicle positions no more than 90 seconds old. Alerts may be up to 10 minutes old. Consumers usually poll on a 15 to 60 second cycle and use conditional requests so unchanged data costs almost nothing to serve.
Why do real-time transit apps show wrong times?
Most often the feed is fine and the app is at fault: a cache that updates every few minutes, a vendor endpoint with its own delay, or a second app that is not reading the agency feed at all. Genuine feed problems come from stale vehicle positions, trips that no longer match the published schedule, or alerts that were never created. Comparing two apps on the same route usually reveals which side is lagging.
Do I need an API key to access a GTFS-Realtime feed?
Agency-published GTFS-Realtime feeds are public and normally need no key. Keys usually appear only with third-party gateways and commercial APIs, where quotas and licensing terms apply. A key embedded in a mobile app can be extracted, so call keyed services through your own backend instead. Also check whether the agency restricts redistribution of its data before shipping anything.
What is the difference between TripUpdate and VehiclePosition?
A VehiclePosition answers where is the vehicle now, carrying coordinates, bearing, speed and sometimes occupancy. A TripUpdate answers when will the vehicle reach a stop, carrying a list of stop-time updates keyed by stop sequence, plus delay, cancellation and schedule relationship values. Both reference the same trip in the static GTFS, and good apps render both: the map for the first, the arrival list for the second.
Conclusion
Real-time transit data feeds run on one short chain: onboard location hardware, an agency server, a dispatch system that joins positions to the schedule, a prediction step, and a protobuf feed that consumer backends poll every 15 to 60 seconds. Everything a rider sees on a phone comes out of that chain, and every failure they complain about traces back to one link in it.
If you are building on real-time transit data feeds yourself, start by picking one agency and one route, pulling the feed and the static GTFS side by side, and looking at how the trip ids line up. That single exercise exposes trip-matching problems, missing positions and stale records faster than any amount of reading.


