During a planned network migration on 2026-08-29 17:35 UTC, we disconnected existing PrivateLink consumer VPC endpoints as part of the cutover. Hybrid customer agents connecting via PrivateLink lost their connections and were not able to reconnect until the new endpoints were provisioned. Full recovery at 18:05 UTC, ~30 min.
What went well
No data loss; agents automatically re-attached with no customer action required.
What could have gone better
Provisioning new endpoints took longer than expected.
We did not proactively communicate to affected customers ahead of the maintenance window.
Follow-ups
We will update our standard procedures for planned maintenance to ensure communication of operations that may cause delays in customer pipelines.
Posted Aug 31, 2026 - 18:17 UTC
Resolved
Impact: Dagster+ hybrid agents connecting to the US region via AWS PrivateLink experienced connection loss. Duration: ~30 minutes (17:35 – 18:05 UTC / 13:35 – 14:05 ET on 2026-08-29)
Summary:
As part of a planned infrastructure migration, we rotated the network load balancers behind our PrivateLink endpoint service. This involved disconnecting the previous set of consumer VPC endpoints so that connections would be reopened on the new infrastructure. During the ~30 minute window between endpoint disconnection and full re-establishment, hybrid agents connecting via PrivateLink were unable to reach the Dagster+ control plane and appeared as unhealthy in our health dashboard.
Once new endpoints came online, agents automatically reconnected and returned to a healthy state without customer action.
What was affected:
Hybrid agents connecting to Dagster+ US region via AWS PrivateLink Job runs, code location updates, and metadata operations dependent on that connection were paused (queued locally) until the connection recovered
What was NOT affected:
The Dagster+ UI and API remained fully available Non-PrivateLink hybrid agents (using public network connectivity) Dagster+ Serverless (EU and US) Dagster+ EU hybrid customers
Timeline (EDT): 12:35 — Existing PrivateLink consumer endpoints intentionally disconnected as part of the migration 13:05 — Full recovery to pre-migration baseline
Root Cause:
Though this was a planned migration, provisioning the new endpoints took significantly longer than expected due to transient errors, leading to noticeable customer impact. This possibility should have been accounted for and communicated to customers in advance.