Senior Site Reliability Engineer
Aviso de fuente externaen Pyramid Consulting, Inc
Skillset:Python, SQL, Incident Response, Root Cause Analysis (RCA), APIs, Kubernetes, GCP, Observability, Monitoring, Logging, Distributed Tracing, SRE, Reliability Engineering, Productio...
Salario
No especificado
Ubicación
Mexico City, Mexico
Tipo de empleo
Tiempo completo
Modalidad
No especificado
Senior Site Reliability Engineer
Mexico City, Mexico
Descripción del empleo
Skillset:Python, SQL, Incident Response, Root Cause Analysis (RCA), APIs, Kubernetes, GCP, Observability, Monitoring, Logging, Distributed Tracing, SRE, Reliability Engineering, Production Support.Location: 100% REMOTE Site Reliability Engineer (SRE) – Incident Response & Reliability EngineeringSite Reliability Engineer (SRE) to improve the stability, reliability, and performance of large-scale production systems. This role combines incident response, automation, observability, and data analysis to identify reliability trends, reduce operational risk, and drive continuous improvement across cloud-native environments.Key ResponsibilitiesRespond to and manage production incidents, perform root cause analysis, and drive preventative solutions.Develop automation and operational tooling using Python and APIs.Analyze incident, log, and performance data using SQL and statistical methods to identify reliability improvements.Build and enhance monitoring, alerting, logging, and distributed tracing capabilities.Support and optimize Kubernetes-based applications running in GCP.Create dashboards, reports, and reliability metrics to communicate technical and business impact.Partner with engineering teams to improve system resilience, scalability, and operational efficiency.Required Skills4+ years of experience in SRE, DevOps, Systems Engineering, or Production Support.Strong Python scripting and automation experience.Advanced SQL and data analysis skills.Experience with incident response, RCA, and production troubleshooting.Strong understanding of observability, monitoring, logging, and distributed tracing.Experience with Kubernetes and Google Cloud Platform (GCP).Experience integrating and working with APIs.Excellent communication skills with the ability to explain technical issues to non-technical audiences.Preferred SkillsExperience with statistical analysis, anomaly detection, or performance trend analysis.Familiarity with Datadog, Splunk, Grafana, Prometheus, or OpenTelemetry.Knowledge of SRE principles, SLIs/SLOs, and resilience engineering.
¿Es tuya esta vacante?
Reclámala gratis y recibe candidatos con video en CazVid.