What GTFS Data Is and How to Use It: A Practical Guide (2026)

GTFS data is the schedule a transit agency publishes so anyone can read it: stops, routes, trips, timetables and fare rules packed as plain-text CSV tables inside one zip file. GTFS stands for General Transit Feed Specification, and because every agency publishes in the same shape, one parser can read the buses in Lisbon and the trains in Osaka.

It describes service as planned, not as running. That distinction is the one that catches almost every newcomer, and it shapes how you should build on top of a feed.

If you build civic apps and urban data tools, GTFS is the most reusable public dataset your city already publishes. The rest of this guide covers what sits inside the archive, where to download it, and how to turn it into something a rider could actually use.

Table of Contents

What Is GTFS Data?

GTFS data is an open data standard for publishing public transport information in a fixed, machine-readable structure. A transit agency exports its schedule from its own operations software into GTFS format, and any app, map or analysis tool can then read that schedule the same way without a custom integration for every city.

The purpose is interoperability. Before GTFS, every agency structured its data differently, so a journey planner covering two cities meant two bespoke importers. One shared format means a single parser gets you route maps, arrival predictions and coverage statistics for thousands of cities.

The canonical specification lives at gtfs.org, run by MobilityData. The Schedule Reference document there was revised on 27 April 2026, and that page is worth keeping open in a tab while you build, because it settles most field-level arguments faster than any forum thread.

Two definitions settle nearly every beginner question:

  • GTFS (General Transit Feed Specification) is the open standard, and by itself almost always means the static schedule format.
  • Feed is one agency’s complete GTFS dataset, published as a single zip archive that is replaced wholesale on each refresh.
  • Service day is the scheduling period a trip belongs to. It starts at noon on the previous calendar day and can run past midnight, so it does not line up with a normal day.
  • Timepoint is a stop where the timetable gives an exact arrival and departure time rather than an estimate.

The format is deliberately plain. There is no database, no API layer and no schema migration to negotiate. You unzip a folder of text files, parse the CSVs and join them on shared IDs.

GTFS Data at a Glance

AspectWhat it looks like in practice
SpecificationAn open, community-governed format maintained by MobilityData at gtfs.org
DeliveryA zip archive of comma-separated text files, one entity per file
Core entitiesAgency, routes, trips, stops, stop times, service calendars and shape geometry
Common publishersTransit agencies, city open-data portals, national transport authorities and their contractors
Refresh cadenceDaily for most urban bus feeds; weekly or monthly is still common for small operators
Typical usesDeparture boards, journey planners, stop finders, network maps, accessibility audits, coverage analysis
Live dataNot included. Vehicle positions and delays arrive separately through GTFS-RT

One more thing to note before you go further: the archive replaces itself entirely. A feed URL does not version releases, so if you cache a download for offline use, cache it with the date you fetched it.

What Files and Fields Are Included in GTFS?

What Files and Fields Are Included in GTFS?

Every feed has the same core files, each one a CSV with a header row. Required files must be present; optional files unlock richer features when an agency includes them.

FilePresenceWhat it describes
agency.txtRequiredThe operator: name, timezone, language, URL, contact details
stops.txtRequiredEvery stop and station, with stop_lat, stop_lon and parent stations
routes.txtRequiredLine-level definitions: route_short_name, route_long_name, route_type, route_color
trips.txtRequiredOne scheduled journey per row, linking a route to a service and a shape
stop_times.txtRequiredArrival and departure times for each stop on each trip, ordered by stop_sequence
calendar.txtRequiredWhich service_ids run on which days of the week, with start and end dates
calendar_dates.txtConditionally requiredExceptions: added and removed service dates that override calendar.txt
frequencies.txtConditionally requiredHeadway-based service instead of a fixed timetable, common for rural and demand-responsive routes
shapes.txtOptionalOrdered latitude and longitude points that draw the route on a map, ordered by shape_pt_sequence
feed_info.txtOptionalPublisher, feed version, feed validity start and end dates, language
transfers.txtOptionalTransfer rules between stops, with minimum connection times
pathways.txt, levels.txtOptionalIndoor navigation inside large stations, including lifts and escalators
translations.txtOptionalStop and route names in other languages
fare_media.txt, fare_products.txt, fare_leg_rules.txtOptionalThe modern fare model, replacing the older fare_attributes.txt approach
attributions.txtOptionalNamed attribution the publisher requires you to display

Stops, routes, trips and stop times are joined by identifier, not by name. Take a bus route 44 with route_id R44. trips.txt holds three rows with different trip_ids on a weekday, one per direction and time band. Each row points to the same route_id, a service_id that says which days it runs, and a shape_id. stop_times.txt then lists that trip’s stops in order, carrying arrival and departure times and a stop_sequence that starts at 1.

Join those four tables and you have a complete journey: where the bus goes, when it leaves, which stops it serves and how the line is drawn on a map. Nothing else in the archive is required to show that itinerary.

How Do Agencies, Routes, Trips, Stops, and Times Connect?

The identifier chain is short enough to memorise: agency to route, route to trip, trip to stop times, stop times to stops, and trip to shape.

Start at a route and you join route_id into trips.txt. From there trip_id joins into stop_times.txt, which holds stop_id, arrival_time, departure_time, stop_sequence and the timepoint flag. stop_id joins into stops.txt for the name and coordinates, and shape_id joins into shapes.txt for the drawn line.

Start at a rider’s request instead and the chain runs the other way. They ask for the next departure from Market Street. That resolves one stop_id, you filter stop_times to future times after converting them to real timestamps, then join to trips to learn the service_id and route_id, and back to stops and routes for the display name.

A few fields cause confusion because the documentation is thin on them. service_id groups trips that share the same calendar pattern, so “weekday morning” and “evenings after 7pm” are often two different service_ids rather than one. block_id groups consecutive trips worked by the same vehicle, which is how you count buses needed for a shift. direction_id is usually 0 or 1 and sometimes reversed between agencies. route_type is an integer from the extended route types list, where 3 means bus and 2 means rail, but older feeds still use 0 and 1.

That is enough structure to reconstruct any itinerary a rider could be shown, and it is the whole of what the static format can tell you.

What GTFS Data Can You Build With?

A departure board for a stop is the smallest useful project: next ten scheduled arrivals with route name and headsign. A stop finder is next, since you already have coordinates and parent stations in stops.txt.

Journey planning needs more work, because GTFS has no routing engine and no transfer graph. You build a time-dependent search over consecutive stop_times rows, then apply transfers.txt rules where they exist. That is a weekend project, not a morning one.

Service calendars and holidays come from calendar.txt and calendar_dates.txt, so you can show “no service on Thanksgiving” without guessing. Accessibility work uses wheelchair_boarding on stop_times and paths.txt, which tells you whether boarding is possible at a given stop, on a given trip, at a given time.

For analysis rather than rider-facing output, GTFS data supports network maps, service frequency plots, coverage and accessibility audits, and commute or site-selection studies that compare transit access across candidate locations. Analysts treat these feeds as the public’s most detailed description of a transport network.

How to Download and Inspect GTFS Data

There is no single global publisher, so finding the feed is part of the work. The Mobility Database catalogues public feeds by country and region and is the fastest first stop. Transitland indexes feeds worldwide and adds editorial metadata such as licensing and modes. Many city open-data portals publish the archive directly, and large agencies usually host a developer page with a feed URL and a data dictionary.

Search terms that work: the agency name plus GTFS, the city name plus open data, and in the US, the state transit authority. Once you have a URL, cache the archive with its fetch date and check the HTTP response before parsing.

curl -L --fail -o feed.zip "https://example.org/gtfs/gtfs.zip"
unzip -o feed.zip -d feed
ls -lh feed
file feed/stops.txt

Inspect before you load. Check which files exist against the required list, confirm encoding is UTF-8 and look for a byte order mark, note the row count of stop_times.txt since it dominates archive size, and read feed_info.txt for the feed version and validity dates.

Check identifiers too. Duplicate stop_id or trip_id values are a common real-world defect, and they will surface as duplicated rows in your app rather than as an import error. Cross-platform case sensitivity bites too, since file names are defined in mixed case and a macOS workspace will happily let you reference stops.txt when the archive holds Stops.txt.

How to Use GTFS Data to Build a Transit App

How to Use GTFS Data to Build a Transit App

Here is the sequence I would follow, in this order. Each step assumes the previous one produced sane output.

  1. Load the archive. Unzip it and read the five core tables with a CSV parser that keeps every column as text. IDs are strings, and treating them as numbers is the most common early mistake.
  2. Validate relationships. Every route_id in trips.txt should exist in routes.txt, every stop_id in stop_times.txt should exist in stops.txt, and every trip_id should have stop times. Report orphans rather than silently dropping them.
  3. Build stop points. Filter stops.txt to location_type 0 or 1 for map display, then build point geometries from stop_lat and stop_lon. Parent stations arrive as separate rows with location_type 1, which is what you want for interchange labelling.
  4. Build route lines. Order shapes.txt by shape_pt_sequence and stitch the points into one line per shape_id. This is a one-liner in PostGIS:
SELECT shape_id,
       ST_MakeLine(ARRAY_AGG(
         ST_MakePoint(shape_pt_lon, shape_pt_lat)
         ORDER BY shape_pt_sequence
       )) AS line
FROM shapes
GROUP BY shape_id;
  1. Join the times. Convert arrival_time and departure_time from HH:MM:SS into seconds past midnight of the service day. Times above 24:00:00 are normal: a 00:31 departure belongs to the previous service day, so 24:31 is the correct value and 00:31 is the wrong one.
  2. Handle cancellations and duplicates. Apply calendar_dates.txt exceptions before you join to calendar.txt, and drop rows whose service was cancelled on that date. Then attach the service date to each time so you have a real timestamp.
  3. Present departures. Pick a stop_id, filter stop_times to future timestamps, join to trips for service and route, then to routes for the display name and headsign.

Steps 5 to 7 in Python look like this:

import pandas as pd

d = "feed/"
stops     = pd.read_csv(d + "stops.txt", dtype=str)
routes    = pd.read_csv(d + "routes.txt", dtype=str)
trips     = pd.read_csv(d + "trips.txt", dtype=str)
stop_times = pd.read_csv(d + "stop_times.txt", dtype=str)

def gtfs_seconds(t):
    h, m, s = (int(x) for x in t.split(":"))
    return h * 3600 + m * 60 + s

# how many trips does route 44 run on a weekday?
tue = trips[trips.service_id.isin(
    pd.read_csv(d + "calendar.txt", dtype=str)
      .query("monday == '1' and tuesday == '1'").service_id)]
print(tue[tue.route_id == "R44"].trip_id.nunique())

# next departures from one stop
board = (stop_times[stop_times.stop_id == "STOP_1021"]
         .assign(t=lambda d: d.arrival_time.map(gtfs_seconds))
         .sort_values("t")
         .merge(trips[["trip_id", "route_id"]], on="trip_id")
         .merge(routes[["route_id", "route_short_name", "trip_headsign"
                        if "trip_headsign" in routes.columns else "route_long_name"]],
                on="route_id"))
print(board.head(10))

Read that snippet as a sketch rather than production code. The mixed-column reference in the last merge is clumsy, and a real app needs the service date applied before comparing times to now. The point is that the heavy lifting is joins, not clever parsing.

How Do You Validate a GTFS Feed?

Run validation before you build, not after something looks wrong on screen. MobilityData publishes a canonical validator, and the reference implementation is maintained as the gtfs-validator project, so you can run the same checks an agency runs on itself.

The checks worth scripting yourself, at minimum:

  • Every required file is present, named with the exact expected case, and parses as CSV.
  • Every foreign key resolves. Orphan route_ids, trip_ids, stop_ids and service_ids are the most common hard failure.
  • IDs are unique within their file, and no row has more columns than the header.
  • Times match HH:MM:SS and are non-decreasing along each trip, allowing for values above 24:00:00.
  • stop_sequence increases within a trip with no gaps or duplicates.
  • Coordinates are in range, with stop_lat between -90 and 90 and stop_lon between -180 and 180.
  • Service dates have not expired. Plenty of feeds are still live while their calendar end date sits in the past, which produces an empty departures board and no error message.
  • feed_info.txt validity dates match calendar.txt start and end dates.

Beyond the file, check the agency-specific requirements listed with the feed. Some agencies publish several variants of the same archive, one trimmed for a consumer app and a full one for mapping, and they are not interchangeable.

What Is the Difference Between GTFS and GTFS-RT?

GTFS Schedule is what agencies publish as a file. GTFS-RT is what they stream over HTTP, and it is a different specification built on Protocol Buffers with three message types: trip updates for delays and cancellations, vehicle positions for where buses are right now, and alerts for unplanned service changes.

GTFS ScheduleGTFS-RT
FormatCSV files in a zip archiveProtocol Buffers messages over HTTP
ContentsPlanned routes, stops, timetables, faresLive positions, delays, cancellations, alerts
DeliveryFile you download or fetchFeed endpoint you poll or subscribe to
RefreshDaily, weekly or monthlyEvery few seconds
ReliabilityStable and completePartial, and many agencies expose only some message types
Typical useTimetables, maps, planning, analysisLive arrival boards and delay alerts

The two work together. A live feed references scheduled trips by trip_id, so if your scheduled data is stale or your IDs have drifted, the realtime data has nothing to attach to and the delays never show up. That stale-schedule problem is the most common cause of “the app says the bus is on time when it is not”, and consumers frequently blame the app rather than the agency.

So never present GTFS Schedule as guaranteed live service. Label it as scheduled times, and add GTFS-RT as a clearly separate layer when the agency provides one.

What Are the Common GTFS Data Challenges?

Licensing and attribution catch people building commercial civic apps. The specification is open, but the data is not automatically free of conditions. feed_info.txt and attributions.txt carry the publisher’s name and any attribution the agency requires, and some cities publish feeds only for non-commercial use or with redistribution restrictions. Read the terms on the download page, not just the file format.

Feeds are operational documents, not perfect documentation. They are generated from scheduling software that people configure by hand, so inconsistent route_type values, swapped direction_id conventions, duplicated stop_ids and mid-feed field changes all show up in the wild. Independent surveys of hundreds of US feeds have found widespread errors, so treat any assumption about a specific feed as something to verify.

Identifiers churn. Agencies renumber stops and routes when they reorganise networks, and a hard-coded stop_id in your app breaks without warning. Store your own mapping layer and treat upstream IDs as replaceable.

Calendar complexity is the quiet one. Trips belong to service_ids, which belong to a weekly pattern plus date exceptions, and a calendar join that forgets exceptions will show you service running on holidays and none on a school in-service day. Timezones add another layer, since a multi-region agency has one feed spanning several zones and a naive parse will silently shift times.

Freshness is uneven. Large urban agencies update daily. Small rural operators often update monthly, or leave the feed running long past its end date. Some feeds redirect, and some are served through a CDN that serves a stale copy for hours after an update. Log the fetch date and the HTTP headers, or you will debug phantom bugs.

Finally, know what GTFS cannot answer. It has no actual running times, no crowding, no ridership, no reliability history and no fares you have to pay. It describes the plan. If your question is about how the service actually performed, you need AVL or smartcard data that the feed does not contain.

Frequently Asked Questions

Is GTFS data real-time?

No. A GTFS Schedule file describes planned service: routes, stops, timetables and fares, published as a zip archive that is usually refreshed daily. Live vehicle positions, delays and cancellations come from GTFS-RT, a separate protobuf-based specification that agencies stream over HTTP. You can build a working timetable app from GTFS alone, but label the output as scheduled rather than live.

Is GTFS data free to use?

The specification is free and open, and almost every agency publishes its feed at no charge. The data carries its own terms, though. Many cities publish under open licences, while others restrict commercial reuse or require visible attribution through attributions.txt or feed_info.txt. Check the download page terms for the specific agency before you put a feed inside a product.

What format is a GTFS file?

A zip archive containing comma-separated text files, one per entity, each with a header row: agency.txt, stops.txt, routes.txt, trips.txt, stop_times.txt, calendar.txt and others. Files are UTF-8 encoded, values may be quoted when they contain commas, and times use HH:MM:SS and can exceed 24:00:00 to represent trips running past midnight.

Do all transit agencies provide complete GTFS data?

No. Coverage varies widely. Major urban agencies usually publish a full feed with shapes, transfers and fares, while small operators may publish only the required files, refresh monthly, or not publish at all. Optional files such as shapes.txt and fare_products.txt are frequently missing, so every feed needs an inspection pass before you rely on a feature.

Can I use GTFS data in a commercial app?

Often yes, but it depends on the agency. Some feeds are published under open licences such as Creative Commons that permit commercial use with attribution, others restrict reuse to non-commercial projects, and a few allow personal and research use only. Read feed_info.txt and attributions.txt, then check the terms on the publisher’s developer page. When you combine feeds across agencies, each set of terms applies separately.

Conclusion

GTFS data is a set of linked plain-text tables describing scheduled transit service, standardised so one piece of code can read a thousand cities. The value is in the joins: routes to trips, trips to times, times to stops.

Start with one feed. Download it, list the files, check the encoding and the feed validity dates, then validate every foreign key before you write a single feature. Once that passes, build the next ten departures from a single stop, and you will understand what GTFS is faster than any definition can tell you.

Leave a Comment