Success Story
FinOps at scale for YALO: ~$2,850/month in confirmed savings and a GKE cluster down from ~165 to ~140 nodes with no hit to stability
YALO is a conversational commerce platform running an intensive mix of microservices on Kubernetes (GKE), databases such as CloudSQL and Redis, Kafka event queues and Datadog observability. Craftech ran an infrastructure assessment and a comprehensive cost-optimization (FinOps) plan on Google Cloud Platform, cutting spend without compromising system stability.
~$2,850/mo
confirmed effective savings (projected up to ~$10,000/mo)
165 → 140
GKE nodes after rightsizing (lows of 134)
−$667/mo
from removing 77 orphaned persistent disks
~$1,000/mo
savings from resizing 6 Redis instances
Challenge
YALO ran a massive production cluster of ~157 nodes on GCP with exceptionally low CPU efficiency, close to 10%, due to widespread over-provisioning. At the same time it faced critical Datadog overages —a projected excess of ~$3,500 USD per month— caused by extremely high-cardinality custom metrics (117 times over the plan limit) and by indexing infrastructure and debug logs that added no value. On the database side, CloudSQL and Memorystore (Redis) instances were paying for enormous capacities (for example 16 vCPUs and 72 GB of RAM) against historical real usage below 15%. And the lack of a tagging scheme made it impossible to slice and analyze costs by end customer or by service.
Solution
Squad 3 executed a comprehensive FinOps plan without compromising stability. It split the ecosystem of more than 120 services into progressive batches for a conservative vertical rightsizing —adjusting CPU/RAM requests and limits to real usage and enabling autoscaling (HPA)— with Python scripts to audit node underutilization. It resized the Memorystore and CloudSQL instances in place after analyzing hit ratio, connections and historical peaks. It contained Datadog costs without touching application code: it applied Metrics without Limits, disabled unnecessary percentiles and created Exclusion Filters to drop debug logs and health-checks. It safely removed dozens of orphaned persistent disks. It implemented KEDA for event-driven scaling (based on Kafka lag), started organizing mixed nodepools with Spot and On-Demand instances, and defined a tagging standard (service, cost-center, customer, tier) for its Pulumi IaC pipelines.
Results
Rightsizing reduced the average number of active GKE nodes from ~165 to ~140 (with lows of 134): Batch 1 alone freed ~80.4 vCPUs and ~87.5 GiB of RAM, Batch 4 freed 30.3 vCPUs and 56.2 GiB, and Batch 7 added another 39.5 vCPUs and 54.7 GiB. In direct GCP savings, resizing 6 Redis instances delivered ~$1,000 USD/month, a single intervention on the main CloudSQL added $700 USD/month, and removing 77 idle disks immediately cut the bill by $667 USD/month. The Datadog reconfiguration projected savings of ~$700 to $1,050 USD/month on custom metrics and $350 to $635 USD/month on logs. The overall impact was ~$2,850 USD/month in confirmed effective savings, with a structured projection of up to $10,000 USD/month once the batch rollout and Spot instances are complete.
Stack & technologies