Case study Python · GeoPandas · MapLibre · Workers Live

How much work is out there?

The attribution work told the owner where his calls came from. The app told him what happened to them. Then he asked something neither could answer: how much roadside work is out there at all, and how much of it does he need to win for the numbers to work? Nobody publishes that. So I estimated it, and built a second tool to find the fleets behind it.

The Corridors tab: a map of Mississippi with every federal-aid highway drawn as a colored line, dark red for corridors carrying 9,000 or more tractor-trailers per day down through orange and pale yellow for under 150. I-20, I-55, and I-10 read as the heaviest lines on the map.
Mississippi ranked by heavy-truck exposure. The dark lines carry the trucks, so they should also carry the blowouts.
54,732→2,747 Raw segments to map sections 10,879 mainline miles, cut where the traffic changes
4.9% Holdout error, down from 18.0% R² 0.833 → 0.989 after one measured correction
156 Lines of server-side code An auth gate. There is no API layer anywhere in the project
§ 01 Origin

A missing denominator

Advertising decisions had gotten better but they'd hit a ceiling. The owner knew what a call cost him and what it earned. He had no idea how many calls were out there to begin with. Was he getting two percent of his market, or forty? Those are different businesses, and they call for different plans.

No dataset counts tire blowouts. There is no authority to look it up from, no trade association publishing it, no way to buy it.

So the tool doesn't try to count them. It models exposure instead: truck miles travelled from federal datasets, on the assumption that blowouts scale with truck-miles travelled. This leaves room for different models based on different blowout rate assumptions with a source of truth to base it all on. That was the question on the table: How much work is out there, and how do we capture it?

§ 02 Corridors

From 54,732 fragments to a ranking

The input is FHWA's HPMS release, the federal highway inventory. It reports traffic counts and vehicle classification per road segment, but it arrives fragmented by arbitrary measurement points, so a single interstate shows up as thousands of disconnected pieces.

A six-stage Python pipeline turns that into something you can look at. Each stage is its own module and runs to a named checkpoint. Before handing off, each one checks its own invariants: mileage conserved, no double inventory, data integrity maintained.

  • Extract — keep the mainline highways and drop everything else. HPMS says what each road is (mainline, ramp, frontage) and how it's signed (Interstate, US, state route).
  • Metrics — truck-miles per segment. Nulls treated as missing data, false zeros located.
  • Corridors — put each road back together in order. I‑20 arrives as 1,822 separate rows, cut at arbitrary measurement points. Sorting by mileposts or counties failed on inconsistencies. What works is the shape itself: match the end coordinate of one piece to the start coordinate of the next and chain them. Once chained, I‑20's mileposts run strictly west to east, which is how I know it worked.
  • Interpolate and estimate — fill short interior gaps, then estimate the rest through a donor chain of past data.
  • Aggregate — cut the chained corridors into sections wherever reported traffic changes by more than the configured tolerance.
  • Export — write the map out four times, once per section size, as four GeoJSON files. Pre-cutting them here means the dropdown just swaps files and the browser never has to re-cut anything.

Granularity that follows the data

This control sets how long each colored piece of road on the map is. It started out as a target: merge neighbors until sections average, say, ten miles. That erases the thing the map is for. Put a 3-mile stretch carrying 4,000 trucks a day next to a 7-mile stretch carrying 200, merge them to hit the target, and the map reports 1,340 across all ten miles. The busy stretch now looks quiet, the quiet stretch looks busy, and the hot spot is gone.

So it splits on traffic instead of length. Wherever volume changes by more than the percentage you pick, the road gets cut there. That 3-mile stretch stays its own section because it differs from its neighbor by far more than 15%. Length stops being a goal and becomes a result: busy areas break into many short sections, uniform rural highway stays in long ones. Moving the slider makes the map coarser or finer without ever averaging away a difference that matters.

A drag-box selection over central Mississippi highlighting 528 road sections in bright blue, with a Selected panel reporting 2,102 miles, 2,191,447 truck-miles, an average of 1,043 tractor-trailers per day, and 34 percent of length measured.
Drag-box selection with running totals: 528 sections, 2,102 miles, 2.19M truck-miles. A service area sized in one gesture. The measured fraction sits right next to the total, so the number carries its own caveat.
§ 03 Honesty

Filling in the gaps

HPMS reports total traffic far more often than it reports vehicle classification. About 45% of corridors have a solid vehicle count and no idea how many of those vehicles were tractor-trailers. Delivering accurate estimates that are honest about their source is vital. The final goal being a complete map with minimal gaps that isn't misleading.

To fill missing gaps, previous years data was first used. Ranging back to 2018, a chronological heirarchy of assigning values was used. For still un-measured segments, only those with measured segments on both sides were interpolated. The rest were left blank, leaving extrapolation to the end user.

Missing data is a fact of life in any dataset. Safe, informed estimates are vital while drawing a line on where the data is and where it isn't.

Historical data was designed to win, and lost

The estimator's donor chain rested on an assumption that sounds obvious. A real historical count of a road should beat a model of that road, so the 2018 National Network readings got priority.

Holdout testing disagreed: 19.8% error for the 2018 anchor against 6.8% for a neighboring 2024 segment. Seven years of drift cost more than being the same piece of road was worth. Demoting the anchor took overall error from 18.0% to 4.9%, and R² from 0.833 to 0.989.

Zeros that weren't zeros

Forty-three segments reported exactly zero tractor-trailers. Take those at face value and they're quiet roads. So I checked them against their neighbors. One stretch reported 20,536 vehicles a day and 1,212 box trucks, then zero tractor-trailers, with segments on both sides of it reporting plenty. That's not a measurement of zero, it's a missing measurement written down as zero. Believing it would have punched artificial holes through real freight corridors.

Estimates never merge into measurements

Estimated values live in their own fields, separate from measured ones, all the way through the pipeline and into the exported GeoJSON. The map renders them in italics with an est tag. Every field has an info button explaining where that number came from, and estimated fields look different than measured ones.

A detail panel for MS Highway 35 showing a 6.4 mile section. Tractor-trailers, single-unit, all trucks and truck-miles per day are in italics tagged est, while all vehicles per day reads 2,816 in plain type. Measured shows 0 percent of length, estimate basis reads nearest measured stretch of this road, and a note explains vehicle class was never sampled here.
MS 35: total traffic is measured, vehicle class never was. The panel says so in plain language instead of showing one confident number.
§ 04 Carriers

Who owns the trucks

Corridors answers volume and where the work is. But if blowout volume isn't enough, how do we hit target profits? The answer is fleet accounts. Fleet accounts are the best work the business takes. They're repeat customers on payment terms that call with regular work. So the question is, how many fleets are near him, and how does he get in touch with them? Up to now, finding them had been word of mouth.

The Carriers tab loads the FMCSA carrier census, filtered to a radius around the business's home base: 17,105 registered carriers inside 75 miles. You can narrow that by operating status (active registration with interstate-authority, active without, revoked, inactive) and by fleet size. Fleet size is the one that matters when you're hunting for somebody running fifty trucks instead of one.

The Carriers tab zoomed out over the Southeast: purple and grey circles sized by carrier count cluster densely around central Mississippi, with a sidebar breaking 17,105 carriers into status groups and fleet-size filters.
17,105 carriers within 75 miles, circles sized by count per ZIP and colored by operating status.
A selected ZIP panel for Brandon 39042 reporting 399 carriers 26 miles from Raleigh, broken into status groups, above a filterable list of carrier names each with a power-unit count: Rankin County School District 349, Brown Bottling Group 82, Killen Contractors 47.
Drilling into one ZIP: 399 carriers, sorted by power units. A prospect list, in the order worth calling.
§ 05 Architecture

Three pieces that never talk at runtime

Python writes GeoJSON into a folder. Wrangler ships that folder to Cloudflare's edge. The browser fetches the files straight off it. There is no API layer anywhere in the project, and the only server-side code is a 156-line auth check borrowed off the business's tires app.

  • Pipeline. Python, GeoPandas and pyogrio for geospatial I/O, pandas, YAML config. Pure batch: static files in, static files out, no database and no server. Tested with pytest.
  • Edge. Cloudflare Workers via Wrangler 4, with static assets bound and run_worker_first so the worker sees every request before the asset server answers.
  • Frontend. One 1,368-line HTML file. No framework, no bundler, no npm dependency, no build step. Plain ES modules with MapLibre GL. Filtering, the color ramps, drag-box select, and per-ZIP aggregation all run in the browser as MapLibre style expressions over the GeoJSON.

All the analysis happens offline, so nothing needs computing when a request comes in.

The gate protects one path

The two tabs need different protection, so they get different protection. Corridor data is public federal highway information without any personal or business information and stays open to anyone. The carrier file holds FMCSA personal data: names, phones, and emails for roughly 17,000 businesses. Though they were gathered from public sources, no consent was obtained from the carriers so the safe move is to keep it private. So the worker guards exactly one path, /carriers.geojson, with borrowed auth from the tires app.

§ 06 Limits

What this doesn't tell you

  • Local roads. County roads and unnumbered local roads aren't in HPMS and aren't on the map.
  • Truck-miles stand in for rolling exposure, not blowouts. Roughness, rutting, speed limit, and urban fraction ride through to the output unmodeled, waiting for a reason to think they'd sharpen the proxy.
  • Call history can't settle it. That history is shaped by current ad coverage, so agreeing with it wouldn't prove anything. That's why it stayed out of scope.

None of this stops the tool from doing its job. Reporting truck-miles and listing nearby fleets by size both survive every limitation on the list. That's why the scope stops where it does.