Site Reliability Engineer

Aviso de fuente externaen TRIMAH TECHNOLOGIES LLC

Design, develop, and implement AI-driven systems and automation tools to enhance the reliability, security, and efficiency of core digital banking platforms. You will own the operational...

Fuente externa - sin verificarhace 5 díasVigente hasta: 6 oct 2026

Salario

No especificado

Ubicación

Columbus, United States

Tipo de empleo

Tiempo completo

Modalidad

No especificado

Site Reliability Engineer

Columbus, United States

Descripción del empleo

Design, develop, and implement AI-driven systems and automation tools to enhance the reliability, security, and efficiency of core digital banking platforms. You will own the operational lifecycle of AI and machine learning models, ensuring high availability, observability, and seamless scaling across cloud environments.
Key Responsibilities:
AI System Design & Implementation: Build and scale automated, AI-driven architectures to optimize digital banking operations and customer touchpoints. SRE & Platform Health: Monitor the health, availability, performance, and SLOs of AI-enabled applications and underlying infrastructure. MLOps & Deployment: Operationalize large language models (LLMs) and machine learning systems into secure, production-grade enterprise settings. Automation & CI/CD: Maintain infrastructure-as-code (IaC) and robust deployment pipelines using modern cloud toolsets. Incident Management: Coordinate with cross-functional teams via enterprise platforms (e.g., ServiceNow) to resolve complex infrastructure and model-serving issues.
Basic Qualifications
Education: Bachelor’s degree in Computer Science, Engineering, Data Science, or a related technical field. Experience: 5+ years of hands-on experience in AI/ML engineering, Site Reliability Engineering (SRE), or DevOps. Core Technical Stack: Proficiency in Python or Java; experience with cloud platforms (AWS, Azure) and containerization tools (Docker, Kubernetes). Operations: Solid grasp of SRE fundamentals (monitoring, alerting, error budgets, SLOs) and CI/CD pipelines. Observability: Familiarity with stack monitoring frameworks (Prometheus, Grafana, ELK) and incident response processes.
Preferred Qualifications :
Direct experience operationalizing LLMs or generative AI in a regulated production environment. Background in MLOps, data engineering, or cloud-native AI pipelines. Strong knowledge of secure coding practices and AI data governance in financial sectors.

¿Es tuya esta vacante?

Reclámala gratis y recibe candidatos con video en CazVid.

Empleos similares