September 28, 2026 · Craftech
How Wiixt brings safety procedures to field workers with a voice AI agent on AWS
Working with Craftech, Wiixt launched a Spanish-language voice agent on AWS so field workers can consult safety procedures and report incidents hands-free.
Field workers in high-risk jobs rarely have their hands free to search a manual. For Wiixt, an occupational health and safety software company in Latin America, that gap meant safety procedures and incident reporting lived on a screen that workers could not always use. Working with Craftech, an AWS Advanced Tier Services Partner, Wiixt launched a Spanish-language voice agent built on Amazon Transcribe, Amazon Bedrock with Claude Haiku 4.5 and Amazon Polly. Workers can now ask about a procedure and hear an answer with its source, or file an incident report by voice.
About Wiixt
Wiixt is a software company whose SaaS platform supports occupational health and safety programs. Its users are workers and supervisors who complete assigned safety courses, consult procedures and report incidents through the platform. Each worker sees only the courses, forms and data enabled for them.
The challenge
The platform had no voice channel. Adding one raised requirements that a standard chatbot does not face:
- Hands-free, low-latency conversation. A worker in the field needs a natural spoken exchange, not a pause of several seconds after every question.
- Answers grounded in the right documents. A safety answer must come from the procedures assigned to that worker, with a reference they can verify.
- Per-user permissions outside the model. The agent must never retrieve content or open a form the worker is not entitled to, regardless of what is asked.
- Reliable actions. An incident report is either filed with a real report number or clearly not filed. The agent cannot claim success it did not achieve.
- Safety before speech. Once a sentence is spoken, the worker has heard it. Any check has to happen before synthesis, inside the latency budget.
The solution: a voice agent with controls outside the prompt
Craftech built the agent on the open-source LiveKit Agents SDK, with AWS services handling each stage of the conversation.
The voice pipeline
Amazon Transcribe streaming converts speech to text. Amazon Bedrock with Claude Haiku 4.5 handles reasoning and decides which tool to call. Amazon Polly speaks the answer back. Voice activity detection, end-of-turn detection and preemptive generation shorten the silence between turns, and Bedrock prompt caching covers the system prompt and tool definitions so they are not reprocessed on every turn.
Managed runtime
The agent runs on Amazon Bedrock AgentCore Runtime, whose built-in JWT authorizer validates each incoming token before the agent executes. The execution role follows the AWS reference pattern, with trust restricted by source account and source ARN.
Authorization in code, not in the prompt
The worker's identity token travels with every backend call. Retrieval is filtered by the courses assigned to that worker, analytics tools are registered only when the worker has data to query, and every form request is checked against the worker's enabled modules before it opens. Conversation state is keyed by user and room, so one session never reuses another's context.
Grounded answers
Hybrid search over a Weaviate vector index returns procedure fragments from the worker's own courses, expanded with neighbouring context. Every procedural answer includes the source page and a link to the document.
Reliable incident reporting
The agent fills the report field by field as the worker speaks, persisting partial state in Amazon DynamoDB. A report counts as filed only when the form is complete and the backend returns a real report number. On any failure, the agent is instructed to say the report was not filed.
Safe analytics
When a worker asks about their data, generated SQL runs against a per-conversation, in-memory SQLite copy, never against the live database. Read-only validation and SQLite's query-only mode add two independent guards.
Safety checks before speech
Amazon Bedrock Guardrails, applied through its standalone evaluation API, checks each transcription before it reaches the model and each response sentence before it is synthesized. The policy covers prompt-attack detection, denied topics, contextual grounding, word filters and PII handling.
The system prompt adds the agent's safety posture: it never authorizes work, always cites its source and escalates severe-risk situations to a supervisor. Tracing runs on Langfuse over OpenTelemetry, alongside Amazon CloudWatch Logs and native Bedrock metrics.
Results
The agent has been in production since January 2026, giving workers a voice channel the platform did not have before.
- A new, hands-free channel. Workers can consult procedures and file incident reports by voice, in Spanish, without navigating a screen.
- Faster turns through measured iteration. Per-stage instrumentation showed model time-to-first-token as the largest share of each turn. Across development iterations in a QA environment, median warm-turn latency dropped from 5.3 s to 2.9 s, roughly 46% faster.
- Answers with a source. Every procedural answer cites the page and document it came from, filtered to the worker's assigned courses.
- Permissions enforced at execution time. Retrieval scope, form access and analytics access are checked in code for every request, independent of what the model decides.
- No false confirmations. A report is announced as filed only when the backend returns a real report number.
- Improvement from real use. When production usage showed the agent offering forms a worker could not file, the team added module validation in code rather than relying on prompt wording.
Lessons learned
- Latency is the constraint once accuracy is solved. Preemptive generation, caching, preloading and model choice each act on a different part of a turn. Only per-stage measurement shows which one to pull.
- Authorization belongs in executable controls. Prompt instructions describe desired behaviour; course filters, module checks and conditional tool registration enforce it.
- Voice has no undo. Whatever is synthesized has been heard, so safety checks must run before speech and their cost must fit the latency budget.
- Explicit acknowledgement beats optimism. Isolated conversation state and a hard rule against claiming success without a report number make failures visible and understandable.
Looking ahead
With the voice channel live, the next steps are measuring latency and cost per conversation against the current production configuration, and expanding end-to-end testing of the deployed system.