Site Reliability Engineer
Aviso de fuente externaen TRIMAH TECHNOLOGIES LLC
Design, develop, and implement AI-driven systems and automation tools to enhance the reliability, security, and efficiency of core digital banking platforms. You will own the operational...
Salario
No especificado
Ubicación
Columbus, United States
Tipo de empleo
Tiempo completo
Modalidad
No especificado
Site Reliability Engineer
Columbus, United States
Descripción del empleo
Design, develop, and implement AI-driven systems and automation tools to enhance the reliability, security, and efficiency of core digital banking platforms. You will own the operational lifecycle of AI and machine learning models, ensuring high availability, observability, and seamless scaling across cloud environments.
Key Responsibilities:
AI System Design & Implementation: Build and scale automated, AI-driven architectures to optimize digital banking operations and customer touchpoints. SRE & Platform Health: Monitor the health, availability, performance, and SLOs of AI-enabled applications and underlying infrastructure. MLOps & Deployment: Operationalize large language models (LLMs) and machine learning systems into secure, production-grade enterprise settings. Automation & CI/CD: Maintain infrastructure-as-code (IaC) and robust deployment pipelines using modern cloud toolsets. Incident Management: Coordinate with cross-functional teams via enterprise platforms (e.g., ServiceNow) to resolve complex infrastructure and model-serving issues.
Basic Qualifications
Education: Bachelor’s degree in Computer Science, Engineering, Data Science, or a related technical field. Experience: 5+ years of hands-on experience in AI/ML engineering, Site Reliability Engineering (SRE), or DevOps. Core Technical Stack: Proficiency in Python or Java; experience with cloud platforms (AWS, Azure) and containerization tools (Docker, Kubernetes). Operations: Solid grasp of SRE fundamentals (monitoring, alerting, error budgets, SLOs) and CI/CD pipelines. Observability: Familiarity with stack monitoring frameworks (Prometheus, Grafana, ELK) and incident response processes.
Preferred Qualifications :
Direct experience operationalizing LLMs or generative AI in a regulated production environment. Background in MLOps, data engineering, or cloud-native AI pipelines. Strong knowledge of secure coding practices and AI data governance in financial sectors.
¿Es tuya esta vacante?
Reclámala gratis y recibe candidatos con video en CazVid.