Success Story

Rebuilding the infrastructure and CI/CD on AWS for Chekin: from daily incidents to occasional ones

Chekin is a Travel Tech platform that automates guest check-in and check-out by integrating with lodging reservation systems, with a presence in more than 20 countries and over 18,000 clients. Craftech rebuilt its infrastructure on AWS, standardized its deployments, and set up observability to stabilize a platform that was failing daily.

Travel Tech / Software Relationship ended with a formal offboarding in March 2023; client inactive
MigracionModernizacion

Challenge

When Craftech first engaged with Chekin, the platform was unstable: multiple services failed daily, the architecture was complex, and monitoring was nearly nonexistent, which made errors hard to find and lengthened incident resolution. Prior deployments did not keep a valid state in the repositories (all images used the latest tag), there were more than 100 services in a single namespace exposed to the internet, manual unversioned processes, and a RabbitMQ deployed inside the cluster that brought down critical services, with no backups and no way to scale.

Solution

Since the client itself did not know the full set of its services or the architecture accumulated over years, Craftech decided to rebuild the infrastructure and migrate it to a known, documented one: it recreated the networking foundations (VPCs, subnets, route tables), implemented new clusters with infrastructure as code, and managed dependencies (relational and non-relational databases, queues). It migrated CI/CD to GitHub Actions and normalized deployments with Helm. To migrate from one cluster to another, it implemented an API Gateway that allowed moving one API at a time with fast rollback, and it set up a custom logging, monitoring, and alerting stack with automatic notifications for both infrastructure and business metrics.

Results

Deployment was automated across all applications and developers' work was simplified, letting them provision new resources far more easily. The number of incidents dropped sharply — from daily and across multiple services to occasional and isolated — and resolution time fell thanks to the new monitoring. Bottlenecks were removed, networks and permissions were segmented, and backups were implemented for critical services with persistent data.

Stack & technologies

Amazon EKSKubernetesTerraformHelmGitHub ActionsAmazon API GatewayAmazon RDSRabbitmq

Facing a similar challenge?

Get a free AWS cost & architecture assessment.

The future of your company
takes off with Craftech

Leverage our AWS expertise to propel your company into the cloud.

Get your free assessment