Director of Cloud Engineering
-
Implemented and managed the rollout of cloud-native AI cost controls to bring Claude Code access to 600+ engineers in 4 days.
Architected centralized API rate limiting, cost governance pipelines, and granular token usage quotas across all engineering business units. Integrated real-time alerts and automated throttle mechanisms to safeguard against anomalous runaway costs while maintaining unhindered developer velocity.
Impact: 600+ engineers enabled in 4 days with zero budget overruns -
Established a new team, centralizing fundamental cloud infrastructure; including accounts/subscriptions, networking, and team IAM.
Formed and scaled the core Cloud Platform Engineering team. Standardized account vending and tenant management using Infrastructure-as-Code (Terraform / AWS Control Tower & Organizations), implementing unified SSO federation, role-based access control, and least-privilege guardrails.
Impact: Reduced account vending from weeks to <30 minutes -
Designed and implemented a global cloud network serving 500+ AWS accounts and 26 regions, centralizing connectivity and accelerating delivery, while also increasing network level observability and security.
Engineered a hub-and-spoke transit network using AWS Transit Gateway, Direct Connect, and cross-region peering topologies. Implemented centralized egress inspection with automated firewall policy distribution, VPC Flow Log analytics, and end-to-end packet observability.
Impact: 500+ AWS accounts connected across 26 global regions -
Worked closely with security teams to establish security baselines across all cloud properties, turning security frameworks (NIST, FedRAMP, CyberEssentials) into baseline config enforcement.
Codified complex compliance controls into automated AWS Config rules, SCP guardrails, and CI/CD policy evaluation using Open Policy Agent. Established automated remediation playbooks to ensure continuous posture compliance across all accounts.
Impact: Automated continuous compliance across NIST, FedRAMP & CyberEssentials -
Developed centralized deployment patterns for security and compliance teams to implement solutions, pushing events to Splunk for triage.
Engineered scalable event-driven streaming pipelines with Amazon EventBridge, Kinesis Firehose, and serverless forwarders to aggregate CloudTrail, GuardDuty, and network flow logs directly into Splunk Enterprise for real-time security operations.
Impact: Sub-second event aggregation for real-time SOC triage -
Formalized FinOps team and practices, classifying and codifying spending reports in the organization. Championed multiple cost-saving initiatives, including adoption of spot instances ($600k monthly), ARM-based instances ($440k monthly), and centralized reservations ($1.4m monthly).
Established an enterprise FinOps operational framework with unit economics, automated anomaly detection, and departmental chargeback reporting. Led engineering adoption of AWS Graviton processors, Spot fleets in container clusters, and unified Savings Plans.
Impact: $2.44M+ recurring monthly savings ($29M+ annualized) -
Pushed multiple re-architecture projects, resulting in $1.4m in yearly savings. One project also saw availability improve from 98.2% to 99.94%.
Spearheaded distributed system redesigns decomposing monolithic single-point-of-failure services into stateless containerized microservices, multi-region database replication, and automated canary routing via Route 53.
Impact: Availability elevated to 99.94% and $1.4M annual cost savings -
Negotiated multi-year $50M+ ARR cloud commitments with multiple cloud vendors.
Collaborated with executive leadership and procurement to model future capacity demands, enterprise discount structures (EDP), and private pricing agreements with major cloud and SaaS providers, maximizing contractual leverage and support tiers.
Impact: Secured enterprise discounts and SLA terms across $50M+ ARR commitments