Operations Infrastructure LeadArchitect
Aviso de fuente externaen Tekshapers
Job Title – Operations & Infrastructure Lead/ArchitectJob Location – Mexico (Remote)Experience level – 15+ YearsJob type – Contract Role OverviewYour primary focus will be split between h...
Salario
No especificado
Ubicación
Mexico City, Mexico
Tipo de empleo
Tiempo completo
Modalidad
No especificado
Operations Infrastructure LeadArchitect
Mexico City, Mexico
Descripción del empleo
Job Title – Operations & Infrastructure Lead/ArchitectJob Location – Mexico (Remote)Experience level – 15+ YearsJob type – Contract
Role OverviewYour primary focus will be split between hands-on operational monitoring and remediation of Borg flex-pool resourcing stockouts across BigQuery environments, and systematically reducing frictions in our CI/CD presubmit pipelines.
Must have:15+ Years of experience, NEED DEVELOPER with Lead/ ArchitectGCPBigQueryPython or GOCI/CD PipelinesKubernnetes
Key Responsibilities1. Capacity Management & Mitigation (70%)Active Monitoring: Track cell-level capacity headroom, identify job sprawl, and monitor for Borg flex-pool resourcing stockouts across BigQuery D3E environments.Incident Remediation: Execute structured, tiered mitigations when capacity stockouts occur.Cell & Resource Optimization: Evaluate Flex ceiling adjustments and perform cell relocations by migrating Playground Systems Under Test (SUTs) or other workloads to alternative cells with verified capacity.Tooling & Configuration: Utilize and maintain Cellmate configurations and Flex API to support new feature requirements and smooth migrations.Escalation Management: Serve as the first line of defense; when standard mitigation tiers are exhausted, manage escalations to SRE or specific owning teams for manual quota and footprint remediation.
2. Presubmit Friction Reduction (30%)Test Health Ownership: Own and improve the test quarantine system to maintain a healthy CI/CD pipeline.Developer Velocity: Actively drive down the volume of slow, flaky, and quarantined tests.Cross-Team Collaboration: Partner directly with bug owners and SWE teams to drive friction resolution, ensuring root causes are addressed promptly.
Qualifications & SkillsMinimum Qualifications:Strong experience with infrastructure operations, site reliability engineering (SRE), or production operations.Hands-on experience managing containerized workloads and cluster management systems (experience with Borg or Kubernetes equivalents is highly desirable).Familiarity with large-scale data warehouse environments (e.g., BigQuery or similar cloud infra).Proven track record of handling real-time capacity constraints, resource quota management, and infrastructure escalations.Solid scripting/automation skills to interact with APIs (e.g., Flex API) and manage configuration tools.
Preferred Qualifications:Direct experience with internal Google infrastructure tools (Borg, Cellmate, Flex API).Strong background in improving developer velocity, specifically managing CI/CD pipelines, test automation, and flaky test mitigation frameworks.Excellent communication and stakeholder management skills to effectively drive bug resolutions with busy development teams.
Impact Statement: This role is critical to the operational health of BigQuery D3E environments. By keeping our infrastructure balanced and our presubmit pipelines green, you will directly unlock engineering velocity for dozens of SWEs.
Tekshapers is an equal opportunity employer and will consider all applications without regards to race, sex, age, color, religion, national origin, veteran status, disability, sexual orientation, gender identity, genetic information or any characteristic protected by law.
¿Es tuya esta vacante?
Reclámala gratis y recibe candidatos con video en CazVid.