Staff AI Engineer

Aviso de fuente externaen Qualitara

Staff AI Engineer (Senior considered) — Conversational Voice AILocation: Latin America (Remote) — full professional proficiency in English required. Company OverviewQualitara is a leading...

Fuente externa - sin verificarhace 9 díasVigente hasta: 12 sep 2026

Salario

No especificado

Ubicación

San José, Costa Rica

Tipo de empleo

Tiempo completo

Modalidad

No especificado

Staff AI Engineer

San José, Costa Rica

Descripción del empleo

Staff AI Engineer (Senior considered) — Conversational Voice AILocation: Latin America (Remote) — full professional proficiency in English required.
Company OverviewQualitara is a leading nearshore software development company dedicated to driving business success through innovative technology solutions. Leveraging top engineering talent across LATAM, we empower our clients by delivering cutting-edge software that meets their strategic objectives.Our client is a market-leading conversation-intelligence and call-management company serving automotive retail. Its platform handles millions of dealership phone interactions, and one of its strategic products is an AI-powered call-training SaaS — salespeople and service advisors practice live, spoken role-play calls against AI "customers," get scored automatically, and managers track improvement. The company is bringing this product in-house onto its own stack ahead of a hard vendor-contract deadline, and the AI engine is the crown jewel of that rebuild.
Role SummaryWe are looking for a Staff (or strong Senior) AI Engineer to own the AI spine of the rebuilt platform — the realtime voice role-play engine and the evaluation system that keeps it trustworthy. This is a hands-on, individual-contributor role embedded in a lean, high-seniority squad. You are the named technical owner of the highest-value, highest-variance part of the build: if the AI engine is right, the product wins; if it isn't, nothing else matters.The core mission: rebuild a production realtime voice agent from first principles (not by reverse-engineering the incumbent), and ship it behind an automated eval/regression harness so quality is provable and survives upstream model drift.
What You Will Work On⭐ Build the realtime voice role-play runtime — computer-audio first (this is what the MVP funds): speech-to-speech via a realtime API (e.g. OpenAI Realtime) or a chained STT→LLM→TTS pipeline; own the latency budget, turn-taking, barge-in, and VAD that make a spoken call feel real⭐ Build the task-evaluation pipeline — a hybrid of deterministic rules and an LLM-as-judge calibrated against human labels, producing per-task pass/fail and an all-or-nothing score managers can trust⭐ Build the eval / regression harness — golden-set scenarios, judge graders, CI quality gates, and weekly drift monitoring (a large share of voice-agent regressions come from upstream model changes, not your code)⭐ Build the persona generator — a config-driven generator (stable per-challenge context + randomized trait layer), with prompts and config versioned and kept separate from code, not thousands of hand-authored profilesOwn speech-to-text integration + speaker timing feeding the transcript, and the CallTrainer summary (deterministic conversation metrics + an LLM qualitative-feedback layer). (Transcript storage and recording lifecycle sit with the full-stack engineers.)Implement guardrails / moderation (separate-model input/output screening; out-of-scope handling) and AI observability (tracing, token/latency/cost)Extend the engine to the telephony path when it's funded — a WebSocket↔SIP bridge (Twilio Media Streams / Pipecat or equivalent) for phone-callback and outbound "mystery shopper" calls. Note: telephony is a later / run-phase capability; computer-audio ships first.Partner with the full-stack engineers who consume your services, and set the AI engineering standards the team builds to
Qualifications⭐ marks the core of this role: candidates must demonstrate all starred items from real, hands-on production experience.⭐ Production realtime / voice AI. You've shipped a live voice agent — speech-to-speech realtime APIs or a real STT→LLM→TTS pipeline — and fought the real problems: latency, turn-taking, barge-in, interruption, audio quality. Chatbot-only LLM experience does not qualify.⭐ LLM evaluation as an engineering discipline. You've built an eval/regression harness — golden sets, LLM-as-judge graders calibrated to human labels (CoT-before-scoring), CI gates, drift detection. This is the moat; "we eyeballed the outputs" does not qualify.⭐ Prompt & persona systems, not prompt-tinkering. Config-driven, versioned prompt/persona generation kept separate from code; structured-output generation with validation; A/B and human-in-the-loop publish gates.⭐ Production LLM application engineering. You've shipped real GenAI features end-to-end: structured-output generation with validation, guardrails / moderation (separate-model screening), AI observability (tracing, token/latency/cost), and cost discipline. Not notebooks or POCs.7+ years building production software, with 2+ years shipping production GenAI/LLM applicationsStrong programming skills in an AI/agent context — Python and/or C# / .NET (the AI/voice-service language is a Sprint-Zero decision; the surrounding platform is .NET/C#/Azure). Familiarity with realtime/agent SDKs (Pipecat, LiveKit, realtime APIs) a plusComfort integrating an AI service into a .NET / C# / Azure product environment and deploying on Azure (Azure OpenAI / Azure AI services a plus)Pragmatic guardrails/safety instincts and cost discipline (LLM + telephony usage is a real budget line)Able to work autonomously, set standards, and be the single named owner of the AI engine
Nice-to-HavesTelephony / audio transport at the protocol level (strongly preferred) — WebSocket↔SIP bridging, media streaming (Twilio Media Streams, Pipecat, LiveKit Agents, or equivalent), connecting a voice agent to a real phone call. It's a confirmed product requirement, but a run-phase capability (computer-audio ships first), so it's not a hard gate for the build hireVector stores / RAG in production (hybrid retrieval, rerank, deterministic citation checks) — a grounded AI assistant is on the roadmapSOC 2 / regulated-environment AI experienceSales-enablement, conversation-intelligence, or training/role-play product domain
What This Role Is NotNot an ML research / model-training role — we are not training or fine-tuning foundation models from scratchNot a "called the OpenAI API once" backend role — realtime voice + eval depth is the barNot a prompt-tinkerer role — prompts here are versioned, evaluated, and engineeredNot a data-science role — this is production AI engineering
Why Join Us?At Qualitara, you will be part of a forward-thinking company that values innovation, quality, and delivering outstanding products to our clients. This engagement places you on a lean, high-seniority embedded squad rebuilding a market-leading AI training platform from first principles, as the named owner of its highest-value, highest-risk component. We offer a competitive salary and opportunities for professional growth in a supportive and flexible remote work environment.

¿Es tuya esta vacante?

Reclámala gratis y recibe candidatos con video en CazVid.

Empleos similares