Senior Site Reliability Engineer (LATAM)
About Reap
Reap is a global financial technology company headquartered in Hong Kong with employees across multiple countries. We enable financial connectivity and access for businesses worldwide by combining traditional finance with stablecoins for efficient money movement.
Through our stablecoin-powered corporate cards, payments, and expense management tools, we streamline financial operations and help businesses scale. Our APIs enable businesses to integrate stablecoin-enabled finance into their own products and services - from issuing Visa cards to facilitating cross-border payments.
Backed by leading investors including Index Ventures and HashKey Capital, Reap is building the future of borderless, stablecoin-enabled finance.
About Reap
Reap is a global financial technology company headquartered in Hong Kong with employees across multiple countries. We enable financial connectivity and access for businesses worldwide by combining traditional finance with stablecoins for efficient money movement.
Through our stablecoin-powered corporate cards, payments, and expense management tools, we streamline financial operations and help businesses scale. Our APIs enable businesses to integrate stablecoin-enabled finance into their own products and services - from issuing Visa cards to facilitating cross-border payments.
Backed by leading investors including Index Ventures and HashKey Capital, Reap is building the future of borderless, stablecoin-enabled finance.
About The Role
Reap is building a Site Reliability Engineering practice, and this role is central to it. We run card issuing, payouts, FX and stablecoin settlement across multiple AWS regions, under PCI DSS and financial regulation. The platform is growing quickly - into new markets, new products, and now agent-initiated payments - and the infrastructure underneath it needs to grow up with it. That means real service ownership, reliability measured in SLIs and SLOs rather than intuition, and a platform that engineering teams can serve themselves from instead of queueing for.
That work is largely still ahead of us, which is the appeal. You will help decide what reliability means at Reap, what the platform looks like, and what good engineering practice is in this domain - rather than inheriting someone else's answers.
This is a deeply hands-on senior individual contributor role. We expect you to lead by example and drive technical excellence through the systems you build, the standards you set, and the way you help the engineers around you level up. The team is deliberately flat and distributed across time zones.
Technologies You'll Use
- Cloud: AWS, multi-account across multiple regions - Transit Gateway, PrivateLink, site-to-site VPN, WAF, KMS
- Compute: ECS/Fargate, Lambda and EKS
- Infrastructure as Code: Terraform and CloudFormation
- CI/CD: GitHub Actions, Argo CD
- Databases: Aurora PostgreSQL, ElastiCache/Redis
- Messaging and streaming: SQS, EventBridge, Kafka
- Observability: New Relic, CloudWatch
- Languages and scripting: Python, Go and Bash
What You'll Do
As a Senior Site Reliability Engineer at Reap, you will join a team of experienced engineers transitioning a DevOps organisation into Site Reliability Engineering. You will own the reliability of a payments platform while rebuilding the foundation it runs on - the lights stay on while the platform gets replaced underneath them. We treat repetitive manual work as a bug in the platform, not a chore for a human, and your job is to delete whole categories of it rather than absorb them faster.
- Define what reliability means here: set SLIs and SLOs with product and engineering teams, introduce error budgets, and make \"ship or stabilise?\" a matter of arithmetic rather than argument
- Take part in our on-call rotation, and help build the incident response and blameless postmortem practice around it
- Drive Reap to full Infrastructure as Code coverage: bring the remaining legacy infrastructure under IaC, and get to no manual provisioning, no drift, and every resource defined, versioned and reproducible
- Consolidate our Terraform estate into a coherent, well-structured codebase with module standards, governance and automated drift detection
- Automate account provisioning and environment setup so that new regions and services can be stood up repeatably and consistently
- Build the golden paths and self-service interfaces that let product and engineering teams provision, deploy and observe without filing a ticket, with escape hatches for the cases they do not cover
- Design and implement an ephemerally environment platform so any developer, and any coding agent, can get an isolated production-like environment on demand and have it cleaned up automatically
- Implement industry-standard observability across logging, metrics and distributed tracing, with alerting that is actionable and trusted
- Own cloud operations for PCI DSS and regulated financial systems - uptime, failover, capacity, disaster recovery and incident response
- Embed security into the infrastructure layer: secrets management, least-privilege IAM, network segmentation and compliance controls
- Build infrastructure that lets AI assistants and agents operate safely: sane blast radius, strong isolation, auditable actions






