September 28, 2026 · Craftech
How Escala rebuilt its multi-tenant AI agent platform on Amazon Bedrock AgentCore
Working with Craftech, Escala replaced its legacy AI agent builder with a multi-tenant agent layer rebuilt end to end on Amazon Bedrock AgentCore Runtime.
Escala's CRM lets commercial teams across Latin America run their sales conversations from a single multichannel inbox. Its AI agents, however, ran on a legacy builder where each agent's entire behaviour lived in one block of text. Working with Craftech, an AWS Advanced Tier Services Partner, Escala replaced that engine with a multi-tenant agent layer rebuilt end to end on Amazon Bedrock AgentCore Runtime. Inference runs on Amazon Bedrock, a managed guardrail checks every turn, and token consumption is metered per tenant. The platform has handled live WhatsApp conversations since August 2026.
About Escala
Escala is a SaaS CRM with a multichannel inbox (WhatsApp, Messenger and Instagram), marketing automation and AI agents. Its tenants are commercial teams across LATAM, including companies in regulated verticals that sign in through federated SAML identity providers.
The challenge
The legacy agent builder defined each agent as a single flat instruction block. That model hit five structural limits:
- Contradictory agents. Branching conversational logic forced into one prompt produced inconsistent or incorrect answers, and it did not scale across tenants with different use cases.
- No auditable safety boundary. A prompt instruction can be circumvented by user input and leaves no evidence that it acted. That kept regulated verticals, such as agents talking about investments, out of reach.
- No memory between conversations. Agents lost all context when a conversation ended.
- Limited retrieval and integrations. Retrieval was vector-only and failed on tabular data, and there was no scalable way to connect tenants' own systems.
- No governance plane. There was no managed guardrail, no per-turn traceability, no quality check before publishing an agent, and observability was manual.
Commercially, lead qualification and follow-up still depended on each tenant's sales team working by hand. Escala needed configurable, safe AI agents as a differentiator for its CRM.
The solution: five architectural decisions
Craftech built Nexo, a runtime that executes tenant-configurable, multi-turn agents, together with a visual builder that acts as its configuration and governance plane. It replaced the legacy builder outright, with no backward compatibility layer and no data migration.
1. Managed agentic compute
The agent runs on Amazon Bedrock AgentCore Runtime and calls Claude Sonnet 5 on Amazon Bedrock through the Converse API, so the model is a configuration value rather than an architectural commitment. The team documented the alternative: the loop could have run on AWS Lambda or Amazon ECS, but sessions, long turns, streaming and per-step traceability are exactly what AgentCore manages. The runtime is never reachable from the internet; its only entry point is a SigV4-signed invocation from the dispatcher.
2. A configurable agent engine
Agents are built from steps with completion criteria and interruption scenarios. On every turn the engine checks which criteria are met, moves focus to the first open step, and lets any scenario interrupt the flow based on the lead's text, a CRM field condition, or both, then resumes where it left off. Agents can write Contact and Opportunity fields, create entities and apply labels in the CRM, send PDF and audio, and hand off to a human or another agent. LlamaIndex Workflows provides a thin tool-use layer, keeping the engine itself fully testable.
3. A decoupled message path
Escala's inbox publishes one event per turn to an Amazon SQS FIFO queue keyed by conversation, with deduplication and a dead-letter queue. A dispatcher AWS Lambda function consumes it, invokes the runtime and hands the reply to Escala's messaging layer. Images and documents are processed with Amazon Bedrock Data Automation before the turn runs, and results are cached so a redelivered message is never processed twice. Conversation state lives in Amazon DynamoDB on demand, with point-in-time recovery; durable per-contact memory stays in the CRM, which the agent reads on every turn.
4. Safety as a managed control
An Amazon Bedrock Guardrail is applied on every turn and returns an explicit intervention flag as audit evidence. Containment is also structural: the agent never sends messages itself, and the tenant decides in the inbox whether responses go straight out or wait for a salesperson's approval. Before any agent can be published, the builder runs a thirteen-check quality validation, and the publication history is immutable by IAM design, with rollback to any prior version.
5. Multi-tenancy and metering
Every log line carries tenant, agent and conversation identifiers, so one search follows a conversation end to end. Token consumption is written idempotently per turn, per tenant, agent and conversation, and exposed through a usage API. Credits are derived at read time from Escala's own equivalence table, so no monetary figure is stored on the platform side. Infrastructure is declared as code with SST v4 on Pulumi and deployed automatically per branch and stage.
Results
The new agent layer reached production on the agreed milestone and now serves live WhatsApp conversations for Escala's tenants.
- On-time production delivery. The backend went live on August 27, 2026 and the frontend on August 28, both through the pipeline, meeting the milestone revised through change management for two customer-requested scope additions.
- Inference on managed infrastructure. Every model call runs on Amazon Bedrock inside Escala's own AWS account, with native guardrails and per-turn traceability the legacy engine lacked.
- Verified safety intervention. The guardrail has been verified to block a request for financial advice and return an intervention flag as audit evidence, the control that makes regulated verticals reachable.
- Quality validation from zero to enforced. Every agent must pass thirteen checks before going live; the reference template and the pilot agent both score 13/13.
- Per-tenant cost attribution. Where the legacy platform counted credits per message, consumption is now recorded per tenant, agent and conversation, split by input, output and cache tokens. Failed writes are queued and replayed idempotently, so no turn goes unmetered.
- Engineering quality gates. 843 backend and 760 frontend tests run on every pull request, with 82% backend statement coverage; no change reaches any environment outside the pipeline.
- No fixed hourly cost. The workload is serverless end to end (AgentCore Runtime, Lambda, DynamoDB on demand, SQS, Amazon S3 and Amazon CloudFront), so cost tracks conversations handled and approaches zero for an idle tenant. Every resource carries nine cost-allocation tags.
Lessons learned
- Test the envelope production actually sends. A required switch from an HTTP API to a REST API in Amazon API Gateway changed how paths and authorizer claims arrived, and hundreds of passing tests missed it. The team added tests built on the real gateway envelope and wrote a postmortem.
- Security has to be operable by the customer. Craftech designed per-repository OIDC roles for CI, then agreed to adopt Escala's existing pipeline workflow under their administration, keeping the OIDC design versioned for a later stage and recording the trade-off in an architecture decision record.
- Integrate against the live service, not the documentation. Several discrepancies surfaced only when testing against the real APIs, and customer dependencies, not internal blockers, were the main schedule risk. A dated checklist separating Craftech's work from items awaiting the customer kept the milestone on track.
Looking ahead
With the runtime in production, Escala can offer tenants configurable agents with auditable safety and per-conversation cost visibility, including in regulated verticals the legacy builder could not serve.