When a routing error looks like a missing route
On 30 September 2026, Railway had a short global routing outage, described in its own incident report. Between roughly 07:35 and 07:40 UTC, new connections to custom domains and domains under *.up.railway.app returned HTTP 404 in every region. Existing connections stayed up. Workloads kept running and no data was lost. Some private networking lookups also failed.
The incident came from two ordinary mistakes stacked together. A deployment ran before its database migration. Then a failed lookup was reported as a valid answer: no route exists. The first mistake caused the failure. The second gave it global reach and hid its meaning.
How the failure travelled
Railway's origin proxies ask a routing service where a request for a domain should go. The service can answer with one route, several routes, or no route when the domain does not exist.
The change behind the incident had two parts. A database schema update modified a data type in the table that stores routing data. It supported an upcoming authentication feature, which was shipped disabled. A new routing service version read the new format. The schema and service were deployed through separate pipelines.
Earlier, the routing service lived in Railway's monolith. One pipeline imposed the required order. Railway later split the service into a distributed control plane to improve resilience during a major cloud outage. After that split, the two pipelines could run in parallel. A race in the configuration allowed the schema update and the service deployment to start together.
The new build finished first. It queried columns that did not yet exist, so every route lookup returned an error. The proxies then treated that error in the same way as a successful lookup with no routes. They returned 404. The same routing service resolves private network DNS, which explains the private lookup failures too.
A fallback could have limited the damage. The proxies had previously been able to use the last route they had seen. This change also altered that behaviour, leaving them unable to fall back. Railway names this as the reason the impact became widespread.
There was no rollback. CI eventually applied the schema update and the next retry succeeded. Railway identified the cause at 07:44 UTC and confirmed full recovery at 07:47 UTC. Its public status post followed at 07:48 UTC, marked as a post facto report.
What engineers should do
Make schema changes safe in either order
A pipeline dependency is useful, but it should not be the only protection. Use the parallel change pattern, often called expand and contract. Add the new column or type beside the old one. Deploy code that can read both formats, and write both where needed. Migrate existing data. Move all readers to the new format. Remove the old form only after compatibility is no longer required.
The test is simple: code and schema must each remain safe whichever one arrives first. This makes a race less interesting. Ordering still matters, but a missed ordering constraint does not immediately become an outage.
Check the schema before accepting traffic
The binary should state which migration version it expects. At startup, or through readiness, it should compare that value with the migration version recorded in the database. If they differ, the instance must remain unready and receive no traffic.
This check belongs in the service, close to the assumption it protects. Railway's remediation follows this principle: each rollout validates the schema version it was authored against, and the blue and green rollout stays blocked until the matching schema exists. Railway also lists serial ordering with deploy order and blast radius checked before merge, plus shard-based control plane rollouts by region.
Keep absence and failure as different types
A route lookup has at least three outcomes, and the interface should preserve them.
- Routes found: forward the request.
- Domain absent: return 404.
- Lookup failed: follow an explicit dependency failure path.
In practical terms, the lookup function should return either a route result, an explicit absent value, or an error. It should never turn an error into an empty route list. For a routing layer, dependency failure should normally serve the last known good routes. If that is impossible, return 503 rather than 404. A 503 tells clients and monitors that the condition is temporary and may be retried. A 404 states that the system knows the route does not exist.
Treat fallback changes as risky changes
Fallback code is quiet until the main path fails. That makes it easy to change and hard to observe. Roll out changes to fallback behaviour separately from schema and service changes. Put them behind a flag. Test them by deliberately making the routing dependency fail, then confirm that known routes still work and unknown routes retain the right meaning.
Combining a new data contract, a new binary and new fallback behaviour creates several ways to fail at once. Separate releases make faults easier to contain and explain.
Write down the guarantees before a split
Moving a service out of a monolith changes more than process boundaries. The old deployment may have supplied implicit guarantees: migration order, atomic release, shared configuration, compatible versions and one rollback point. Once pipelines are separate, those guarantees need explicit owners and enforcement.
Before a split, list every guarantee the monolith currently provides. For each one, decide whether the new design preserves it, replaces it, or deliberately removes it. Put the surviving rules in CI, readiness checks and rollout policy. Architecture diagrams rarely show deployment order, yet production depends on it.
Monitor what users should receive
This outage returned clean HTTP responses with the wrong meaning. A monitor that alerts only on 5xx responses or timeouts could have stayed silent. Requests did not reach customer services, so their internal dashboards would mostly have shown a drop in traffic, as IncidentHub's analysis notes.
External probes should request known public domains and verify the exact expected response, including status and useful content. Alert on sudden traffic drops as well as error rates. Add a health check that resolves and reaches a known service through private networking. Probe frequency matters too. A five-minute interval combined with a two-failure threshold can miss an outage this short.
The lesson
Railway deserves credit for publishing a fast, specific report. The lesson is one line: make deployment order defensible, and never report failure as knowledge.