Staff AI Engineer
External job listingat Qualitara
Staff AI Engineer (Senior considered) — Conversational Voice AILocation: Latin America (Remote) — full professional proficiency in English required. Company OverviewQualitara is a leading...
Salary
Not provided
Location
San José, Costa Rica
Employment type
Full time
Workplace
Not provided
Staff AI Engineer
San José, Costa Rica
Job description
Staff AI Engineer (Senior considered) — Conversational Voice AILocation: Latin America (Remote) — full professional proficiency in English required.
Company OverviewQualitara is a leading nearshore software development company dedicated to driving business success through innovative technology solutions. Leveraging top engineering talent across LATAM, we empower our clients by delivering cutting-edge software that meets their strategic objectives.Our client is a market-leading conversation-intelligence and call-management company serving automotive retail. Its platform handles millions of dealership phone interactions, and one of its strategic products is an AI-powered call-training SaaS — salespeople and service advisors practice live, spoken role-play calls against AI "customers," get scored automatically, and managers track improvement. The company is bringing this product in-house onto its own stack ahead of a hard vendor-contract deadline, and the AI engine is the crown jewel of that rebuild.
Role SummaryWe are looking for a Staff (or strong Senior) AI Engineer to own the AI spine of the rebuilt platform — the realtime voice role-play engine and the evaluation system that keeps it trustworthy. This is a hands-on, individual-contributor role embedded in a lean, high-seniority squad. You are the named technical owner of the highest-value, highest-variance part of the build: if the AI engine is right, the product wins; if it isn't, nothing else matters.The core mission: rebuild a production realtime voice agent from first principles (not by reverse-engineering the incumbent), and ship it behind an automated eval/regression harness so quality is provable and survives upstream model drift.
What You Will Work On⭐ Build the realtime voice role-play runtime — computer-audio first (this is what the MVP funds): speech-to-speech via a realtime API (e.g. OpenAI Realtime) or a chained STT→LLM→TTS pipeline; own the latency budget, turn-taking, barge-in, and VAD that make a spoken call feel real⭐ Build the task-evaluation pipeline — a hybrid of deterministic rules and an LLM-as-judge calibrated against human labels, producing per-task pass/fail and an all-or-nothing score managers can trust⭐ Build the eval / regression harness — golden-set scenarios, judge graders, CI quality gates, and weekly drift monitoring (a large share of voice-agent regressions come from upstream model changes, not your code)⭐ Build the persona generator — a config-driven generator (stable per-challenge context + randomized trait layer), with prompts and config versioned and kept separate from code, not thousands of hand-authored profilesOwn speech-to-text integration + speaker timing feeding the transcript, and the CallTrainer summary (deterministic conversation metrics + an LLM qualitative-feedback layer). (Transcript storage and recording lifecycle sit with the full-stack engineers.)Implement guardrails / moderation (separate-model input/output screening; out-of-scope handling) and AI observability (tracing, token/latency/cost)Extend the engine to the telephony path when it's funded — a WebSocket↔SIP bridge (Twilio Media Streams / Pipecat or equivalent) for phone-callback and outbound "mystery shopper" calls. Note: telephony is a later / run-phase capability; computer-audio ships first.Partner with the full-stack engineers who consume your services, and set the AI engineering standards the team builds to
Qualifications⭐ marks the core of this role: candidates must demonstrate all starred items from real, hands-on production experience.⭐ Production realtime / voice AI. You've shipped a live voice agent — speech-to-speech realtime APIs or a real STT→LLM→TTS pipeline — and fought the real problems: latency, turn-taking, barge-in, interruption, audio quality. Chatbot-only LLM experience does not qualify.⭐ LLM evaluation as an engineering discipline. You've built an eval/regression harness — golden sets, LLM-as-judge graders calibrated to human labels (CoT-before-scoring), CI gates, drift detection. This is the moat; "we eyeballed the outputs" does not qualify.⭐ Prompt & persona systems, not prompt-tinkering. Config-driven, versioned prompt/persona generation kept separate from code; structured-output generation with validation; A/B and human-in-the-loop publish gates.⭐ Production LLM application engineering. You've shipped real GenAI features end-to-end: structured-output generation with validation, guardrails / moderation (separate-model screening), AI observability (tracing, token/latency/cost), and cost discipline. Not notebooks or POCs.7+ years building production software, with 2+ years shipping production GenAI/LLM applicationsStrong programming skills in an AI/agent context — Python and/or C# / .NET (the AI/voice-service language is a Sprint-Zero decision; the surrounding platform is .NET/C#/Azure). Familiarity with realtime/agent SDKs (Pipecat, LiveKit, realtime APIs) a plusComfort integrating an AI service into a .NET / C# / Azure product environment and deploying on Azure (Azure OpenAI / Azure AI services a plus)Pragmatic guardrails/safety instincts and cost discipline (LLM + telephony usage is a real budget line)Able to work autonomously, set standards, and be the single named owner of the AI engine
Nice-to-HavesTelephony / audio transport at the protocol level (strongly preferred) — WebSocket↔SIP bridging, media streaming (Twilio Media Streams, Pipecat, LiveKit Agents, or equivalent), connecting a voice agent to a real phone call. It's a confirmed product requirement, but a run-phase capability (computer-audio ships first), so it's not a hard gate for the build hireVector stores / RAG in production (hybrid retrieval, rerank, deterministic citation checks) — a grounded AI assistant is on the roadmapSOC 2 / regulated-environment AI experienceSales-enablement, conversation-intelligence, or training/role-play product domain
What This Role Is NotNot an ML research / model-training role — we are not training or fine-tuning foundation models from scratchNot a "called the OpenAI API once" backend role — realtime voice + eval depth is the barNot a prompt-tinkerer role — prompts here are versioned, evaluated, and engineeredNot a data-science role — this is production AI engineering
Why Join Us?At Qualitara, you will be part of a forward-thinking company that values innovation, quality, and delivering outstanding products to our clients. This engagement places you on a lean, high-seniority embedded squad rebuilding a market-leading AI training platform from first principles, as the named owner of its highest-value, highest-risk component. We offer a competitive salary and opportunities for professional growth in a supportive and flexible remote work environment.
Is this your job posting?
Claim it for free and receive video applications on CazVid.