One bad monitoring location should not define an outage. PulsePing requires a majority of locations to fail the same interval before marking a test down. pulseping.com
Learn more at pulseping.com
#UptimeMonitoring #SRE #DevOps

One bad monitoring location should not define an outage. PulsePing requires a majority of locations to fail the same interval before marking a test down. pulseping.com
Learn more at pulseping.com
#UptimeMonitoring #SRE #DevOps
AI is changing how teams respond to incidents, but your tabletop exercises probably haven't adapted yet. Here's why that gap matters for readiness.
Silent concerns kill reliability. Engineers who spot problems but stay quiet might be your biggest production risk—here's why psychological safety matters for…
https://www.reddit.com/r/devops/comments/1x0pt2p/the_biggest_production_risk_might_be_the_engineer/
⚡ Remote Senior DevOps Engineer at Nextiva.
✅ Apply here: https://jobicy.com/jobs/152901-senior-devops-engineer-7
Wrong Procedure, Followed Correctly #incidentresponse #postmortem #devops #sre #cloud This is a clip from our recent Ship It Weekly Podcast episode. Visit https://shipitweekly.fm or link in bio to listen to the full episode!
Error trackers excel at grouping crashes, but structured logs tell you when a transaction silently failed. Here's how to use both to debug the…
RabbitMQ's dead-letter queue buttons have hidden gotchas. Learn what Nack, Reject, and Automatic ack actually do—and where they differ between…
What happens when your Kubernetes pod reports Ready but the service is broken? This practitioner deliberately broke an LLM platform to expose…
Failures You Can Never Engineer Away #sre #reliability #devops #cloud #incidentmanagement #aws This is a clip from our recent Ship It Weekly Podcast episode. Visit https://shipitweekly.fm or link in bio to listen to the full episode!
México y EE. UU. destacan avances en cooperación de seguridad
Roberto Velasco y el embajador Ronald Johnson resaltaron los progresos bilaterales en la agenda de seguridad con respeto a las soberanías.
Silent job failures are invisible to health endpoints. Pair external uptime monitoring with internal telemetry using a stable run ID to…
What changed before an application became unhealthy? 👀
On Oct. 15, see how AI agents can use deployment history, cluster health, logs, and recent changes to investigate delivery issues through Akuity.
Register: buff.ly/HqTuOFi
Our disaster recovery plan was stored in the wiki.
The wiki is hosted on the server that just went down.
Technically, we did have a plan.
The latest update for #Komodor includes "Building AI #SRE Agents, Part 3: Autonomous in the #Cloud" and "Context Engineering for AI Agents: What to Feed an Agent and What to Leave Out".
#DevOps #Kubernetes #CloudNative https://opsmtrs.com/33UjvEK
The latest update for #Last9 includes "HAProxy #Monitoring: Stats, #Prometheus Metrics and Alerts" and "ClickHouse Monitoring: Key Metrics, System Tables and Alerts".
#Observability #TimeSeries #SRE #DevOps https://opsmtrs.com/41EhXHh
The latest update for #Cortex includes "The Five Levels of the Enterprise #AI Software Factory" and "What is Spotify Backstage?".
#microservices #SRE #devops https://opsmtrs.com/3U19Lxq
Reliability Systems: Are They Too Complex? Fixing Failures #sre #reliability #devops #cloud #aws This is a clip from our recent Ship It Weekly Podcast episode. Visit https://shipitweekly.fm or link in bio to listen to the full episode!
Almost every day, another AI SRE product appears. The demonstrations all follow one script. A chat box takes a question about an incident, a few tool calls run, and seconds later a nicely formatted response lists probable causes and suggested actions.