After weeks of embarrassing outages, migration deadlines, carrier failures, and machine-traffic spikes, the most compelling service management story right now isn't any single incident. It's how ownership’s evolving.
What kept jumping out at me as I worked through this week’s updates was how often the customer-facing failure sat between systems rather than perfectly inside one of them. That makes this less a list of incidents than a week about who can actually see, own, and explain the whole delivery path.
Customer journeys now cross cloud platforms, carrier networks, third-party APIs and autonomous agents, while status pages and operating teams still tend to stop at organizational boundaries. You can see that in every story in this roundup.
Cloudflare and Fastly are redesigning infrastructure for machine traffic. Twilio's carrier board shows why an operational platform can still sit above a broken customer journey. GitHub is showing what useful incident accountability looks like.
Atlassian is moving Opsgenie customers while changing the Jira Service Management destination underneath them. New Relic, Dynatrace, ServiceNow, NVIDIA, Microsoft, and Atlassian are all, from different angles, turning agent observability and containment into production operations.
For CX leaders, 'is the service up?' is becoming the wrong first question. The more pressing test is whether you can see the full delivery path, work out who owns the break, and explain what changes next.
TL;DR:
-
Machine traffic is now operational traffic. Cloudflare and Fastly are giving agents dedicated controls, APIs, and runtime visibility. Cloudflare’s new Issues feature now pushes production failures and diagnostic context into agent or on-call workflows too.
-
Reliability still breaks at the handoff. Twilio's carrier incidents show why platform uptime doesn't equal journey uptime. Microsoft’s September 30 Azure gateway incident put the same problem into cloud connectivity.
-
Migration and incident transparency matter together. Opsgenie's 2027 deadline arrives while JSM changes, and GitHub's RCAs show what buyers should demand when systems fail.
-
Agent observability is turning into agent control. New Relic and Dynatrace are measuring the problem; ServiceNow and NVIDIA are moving toward containment, while Microsoft and Atlassian are building governance into execution itself.
What Was Announced During Cloudflare Birthday Week 2026?
Cloudflare Birthday Week 2026 spent a lot of time on machines behaving like internet users. In its September 27 founders’ letter, Cloudflare said automated traffic had already overtaken human traffic on its network in May, beating its own late-2027 prediction by well over a year.
“The Internet is changing more today than at any point since Cloudflare launched back on September 27, 2010.”
Cloudflare threw another big number into the mix during the week: daily AI-agent requests were up more than 1,700% in a year. Traffic growing that fast can’t be filed away as a bot-policy issue. It’s something capacity teams, network teams, and service managers have to plan around.
Cloudflare’s also giving those agents more operational reach. Its new cf agentic CLI covers more than 3,000 API operations, letting agents deploy Workers, monitor infrastructure and configure services. Agents had already risen from 25% of Wrangler usage in March to 48% shortly before launch.
There was another announcement in the form of “Issues” too, which groups Worker exceptions and 5xx responses with the telemetry needed to investigate them. Cloudflare says Issues uncovered two previously obscured Workflows bugs within a day of internal use.
That makes Ruth Kirby’s point from one of our webinars especially relevant:
“I can see exactly what’s going on in real time and my information is being updated in real time.”
Cloudflare pushed that idea further on October 2. Traces can now follow an individual request through security rules, caching, routing, Workers, and the origin, giving teams something much closer to the path a customer actually experiences.
There was a connectivity move as well. Cloudflare and Deutsche Telekom are directly interconnecting their networks, with both companies talking up stronger end-to-end resilience.
For me, that’s the useful Birthday Week takeaway: more machine traffic means more places for a customer journey to break, and more pressure to see the whole route when it does.
The Opsgenie Sunset: What Should Customers Do Before April 2027
The Opsgenie sunset is getting harder to treat as a problem for next year. On September 14, Opsgenie and Jira Service Management were hit for around 100 minutes, with alerts delayed by an average of 22 minutes. Atlassian blamed a configuration change that set off redundant calls, rate limiting, and retries until worker threads were exhausted.
Customers have until April 5, 2027, to move. After that, Opsgenie becomes inaccessible and unmigrated data is deleted. Atlassian’s in-app migration tool moves customers into Atlassian Cloud, rather than Jira Service Management Data Center.
There’s plenty of competition for those customers. PagerDuty is pushing more incident work into Slack, while incident.io says changes to Investigations cut the median time to an accurate incident-channel update from 6.7 minutes to three.
That’s why I’d test the ugly part of the migration first. Do alerts still arrive under pressure? Do escalation rules behave properly? Does somebody actually get paged?
Atlassian is already rolling out circuit breakers designed to keep alert creation, notifications, and on-call schedules working when JSM search is unhealthy. RovoOps is also moving into Slack incident and alert channels.
Then came October 1. A JSM incident delayed email-to-ticket processing and notifications, while some new support tickets temporarily lost their SLA component. Atlassian said response times and backlog management could be affected.
At Team ’26 Europe, Atlassian is also previewing Proactive Service Management, including silent fixes, employee-approved remediation and Rovo Desktop.
The destination is getting smarter while customers move toward it. That makes early failure testing more important, not less.
Why Is AI Traffic Growth Becoming a Customer Experience Issue?
Fastly’s latest network data gives the agent-traffic story some useful scale. Machine-generated requests accounted for more than half of traffic on its network in July and August, while AI requests grew around 6.5x faster than human traffic from January to May.
The more revealing number is where those requests go. Fewer than 9% of human requests reach origin infrastructure, compared with more than 51% of automated AI traffic. Agents are looking for live prices, availability and account information, so they’re reaching much deeper into the APIs and systems behind the customer experience.
Fastly’s response includes AI Runtime Control, AI Firewall and expanded API Security. As Chief Product Officer Kelly Shortridge put it:
“To effectively adopt coding agents and implement AI features without losing velocity or eroding existing business resilience, enterprises need control in production, at runtime.”
Cloudflare is seeing the same shift, while Akamai and MuleSoft are extending governance across APIs, agents and MCP servers. The common problem is fairly practical: infrastructure built around people making requests now has software making them continuously.
Fastly made that point more tangible on September 28. Comcast and Fastly are putting Fastly delivery software inside Comcast’s distributed network for high-demand Peacock live events, bringing capacity closer to viewers when millions arrive at once.
That’s why this matters beyond infrastructure teams. Customers increasingly arrive through agents, while the services answering them still depend on APIs, origins, and network paths that can buckle under very different traffic patterns.
Why Do Carrier Problems Break Messaging When Twilio Is Up?
Twilio’s September status board gave us a useful reminder that platform uptime and customer experience aren’t the same thing.
At the start of October, Twilio was carrying AT&T US SMS emergency maintenance, UK carrier maintenance affecting APIs such as SIM Swap and Silent Network Auth, and delayed SMS delivery receipts for CTBC subscribers in Brazil.
Several incidents sat under External Connectivity rather than Twilio Services. That distinction helps operators work out where the fault sits. It means very little to somebody still waiting for a verification code, appointment reminder, or callback.
The same pattern was showing up elsewhere. Sinch was investigating delayed messages and missing delivery receipts on Ooredoo Qatar, while Infobip recorded temporary SMS degradation across Spanish networks. CPaaS platforms still depend on carrier routes, interconnects and regional networks that sit outside the neat boundary of the platform itself.
That makes Twilio’s wider 2026 strategy interesting. The company is putting more emphasis on conversation quality, Branded Calling and performance monitoring. The harder question is how far that visibility can follow the journey once traffic leaves Twilio.
By October 3, the board had moved again, with SMS delays to Banglalink in Bangladesh and delays or failures involving Hi3G Italy. Twilio also closed a separate London interconnection alert after confirming there had been no customer impact.
This is a good reminder that great status reporting isn’t only about flagging problems quickly. It’s also about closing the loop clearly when a worrying signal turns out not to have broken the customer journey.
What Do GitHub’s September Outages Tell Us About a Good Post-Mortem?
GitHub had a rough September, after a tricky August, but at least its incident write-ups gave buyers something to work with.
The September 28 Copilot Code Review failure is a good example. Roughly 70% of requested reviews failed between 19:23 and 22:08 UTC after a dependent-service change stopped supplying data the feature needed.
More importantly, GitHub explained why it didn’t catch the problem sooner. Review completion is delayed by design, so the signal operators were watching was slow to expose the damage. GitHub is now adding faster failure signals, better dependency tracking, end-to-end deployment validation, and deployment-pipeline changes.
I’d give that an A-. It tells customers what broke, why detection lagged, and what changes next.
The September 13 outage was similarly candid. Around 28 services were affected and, at peak, roughly 96% of web issue-creation attempts failed. GitHub traced the problem to an internal cleanup job putting pressure on a shared authorization database, then admitted its safeguard was watching “only one health signal”: replica lag.
A September 20 Pull Requests incident added another example. A maintenance job on an unusually large repository almost exhausted one Git storage server’s memory. GitHub responded by capping per-job memory use and improving monitoring.
GitLab and Atlassian have published similarly specific explanations this month.
A post-mortem needs to say what broke, why the existing safeguards didn’t stop it, and what’s been changed so the same failure is less likely to happen again.
AI Agent Observability Becomes a Service Management Priority
New Relic’s 2026 Observability Forecast brings numbers to the visibility and observability gaps companies are still struggling with. Its survey of 2,575 IT and engineering professionals found 25% of organizations already have production AI agents running without monitoring.
Those surveyed businesses lose an annualized average of around $74 million to high-impact outages. There’s an ROI signal too: 42% of organizations monitoring agents reported 3x or greater observability ROI, compared with 21% among those running agents without agent monitoring.
Jim Young shared a memorable line:
“AI made software faster to build. It also made it harder to run.”
New Relic is trying to cut out the early detective work. Headless observability gives agents context without making them hunt for it, and Compound Alerts bring related conditions together so the first signal has some substance.
Dynatrace has its own angle. AI model monitoring is now a top use case for 67% of SRE respondents, and the Arize deal gives it more AI traces and evaluation alongside production telemetry. Its MCP work then opens that data up to coding assistants too.
Cisco adds an uncomfortable twist. Its testing found agent-run tasks can generate up to 450% more network traffic. So the service management question is now: when an agent “fixes” something, can you see exactly what it touched, what changed, and whether the incident actually got better?
How Do You Stop an AI Agent Without Making Things Worse?
Spotting an agent that’s behaving badly is only part of the service-management problem. The tougher test is whether teams can contain it quickly without triggering another outage in the process.
ServiceNow’s September AI Control Tower update gets unusually close to that problem. Its kill switch can revoke credentials across ServiceNow agents, Okta, Google Cloud, and AWS Bedrock, with ServiceNow saying containment can fall from around 30 minutes to seconds. The important bit is that the action is logged and reversible.
The same thinking now extends to MCP. ServiceNow’s AI Gateway sits between agents and MCP servers, checking policy when a tool is called and giving teams a clearer view of what is actually running.
Version 3.4 makes that more precise. Teams can pause one MCP server without pulling every other server offline, then restore it later with its configuration intact. That may sound like a small product detail, but it gets at a very real operational problem: the emergency brake should not become the next incident.
NVIDIA is tackling the same issue lower down the stack. Its Open Agent Safety Platform combines OpenShell runtime boundaries with Sentry, a watchdog running on BlueField-4 hardware. NVIDIA says Sentry can quarantine an agent in milliseconds if it steps outside policy.
Microsoft is leaning on isolation instead. More than one million Container Apps Sandboxes run each day across services including GitHub Copilot, Copilot Studio and Azure SRE Agent.
Ultimately, agent operations are starting to look like operations, full stop. Observability tells teams what happened. Service management now has to decide who can intervene and how safely they can do it.
In Other News
A couple of other news stories worth watching from In Beat this week:
-
Azure gateway failure: Microsoft says its September 30 incident began when a gateway-management change collided with separate OS servicing across several regions. The interaction created excess load and stopped affected gateway services scaling properly. It’s a useful reminder that failures often emerge between systems that looked healthy on their own.
-
Cloudflare adds post-quantum visibility: Cloudflare now exposes per-connection post-quantum telemetry through Traffic Analytics, Log Explorer and Logpush. Around 70% of browser traffic reaching Cloudflare uses hybrid ML-KEM, compared with roughly 15% of connections from Cloudflare to origins. You can’t manage a migration gap you can’t see.
-
DigitalOcean tests its redundancy: On October 3, multiple upstream paths in BLR1 failed, including redundant links. Traffic shifted onto a remaining connection that lacked enough capacity and saturated. Managed Databases, Functions, Load Balancers, Spaces and Monitoring were among the affected services.
What to watch: Atlassian Team ’26 Europe, October 6–8, is the one I’d watch most closely for this beat. Atlassian has already scheduled a service-management solution keynote and a JSM first look that promises proactive detection and remediation before employee impact. WebexOne runs October 5–8 too, with Cisco explicitly tying its Intelligent Workplace Experiences story to AI Collaboration, Secure Networking and Smart Spaces. Both could move this roundup’s argument from observing failures toward preventing or containing them.
If you have any recent news about service management and connectivity, get in touch.
Service Management News: Reliability Is Becoming an Ownership Problem
The nexus running through this week’s service management news is responsibility.
Fastly is watching machines put fresh pressure on origin systems. GitHub is tracing failures through shared databases. Twilio's carrier board keeps exposing the handoffs customers never see. Atlassian is asking customers to move deeper into JSM while the Opsgenie sunset clock keeps ticking. Meanwhile, New Relic, ServiceNow, and NVIDIA are all treating agent monitoring and containment as production operations rather than an AI side project.
For CX leaders, that creates a harder question than “is the service up?”
Who owns the customer experience when the failure crosses from one system, supplier, or automated agent into another?
That's where incident management is heading next. The strongest operators will be the ones who can explain what happened to the customer, where responsibility changed hands, when an automated actor needs to be contained, and what they're doing about it.
FAQs
What did Cloudflare announce during Cloudflare Birthday Week 2026?
Confirmed releases included cf, an agentic CLI spanning 3,000+ Cloudflare API operations, Forge, BEACON and Issues, which can route Worker error context into coding agents or on-call workflows. Cloudflare launched end-to-end Traces and a broader observability layer on October 2. It said automated traffic overtook human traffic in May 2026, daily AI-agent requests grew more than 1,700% year over year.
When does the Opsgenie sunset happen?
April 5, 2027 is the cutoff. After that, Opsgenie won’t be accessible and any customer data left behind will be deleted. Teams need enough runway to move alerting and on-call work, then actually test it. Integrations, APIs, escalation rules, and incident communications are exactly the bits that shouldn’t be left until the deadline is looming.
Why is AI traffic growth becoming a CX issue?
According to Fastly, machine-generated traffic became the majority on its network in July 2026, and AI requests grew roughly 6.5 times faster than human traffic from January through May. There’s another wrinkle: more than 51% of automated AI requests make it all the way to origin infrastructure, versus under 9% of human requests.
What makes a good incident report?
A useful incident report explains customer impact, the failure mechanism, why controls missed it, and what changes next. GitHub’s September 13 and 20 RCAs did that, and its September 28 Copilot Code Review write-up now does too: roughly 70% of requested reviews failed, GitHub named the dependent-service change and delayed detection signal, then listed the corrective work.