Senior Site Reliability Engineer (Performance and Scalability)
Role overview
Your mission is to make DigitalZone able to scale. You will build the platform's capacity to absorb campaign-level traffic spikes, and you will give every engineering team the tools, standards, and practices to load- and failure test their own systems. This is an enablement role at its core: you raise the reliability bar across the org by building capability, not by owning every service yourself. What you'll do • Build the platform's scalability foundation: capacity planning, autoscaling, caching, queueing, and graceful degradation designed for large campaign spikes rather than steady-state load. • Establish load and failure testing as a standard engineering practice, giving teams the frameworks, tooling, and runbooks to test their own services and act on the results. • Own SLOs, error budgets, and the observability stack (metrics, logs, traces, alerting) across TypeScript, Go, and PHP/Laravel services, and standardize how teams instrument for scale. • Harden Postgres and AWS infrastructure for performance and availability, and reduce toil through automation and IaC. • Lead incident response and blameless postmortems, and drive the systemic fixes upstream into design and campaign planning so reliability is built in, not bolted on. • Partner with engineering teams early on capacity and resilience, acting as the multiplier that makes them self-sufficient at scaling their own systems.
Requirements What you'll bring • 5+ years in SRE, platform, or backend engineering, with strong production ownership of large-scale systems operating at 10s of thousands of requests per minute. • A track record of scaling systems through real traffic spikes, and of designing and running load and failure testing programs that other teams adopted. • Deep AWS experience and a solid grasp of Postgres performance and scaling. • Fluency with observability tooling and infrastructure-as-code, plus scripting in Go, TypeScript, or similar. • A calm, systematic approach to incidents, and the communication skills to influence and enable other teams rather than gatekeep.
Benefits • Immediate, large-scale impact on a high-growth business • Top-of-the-market compensation packages • Work alongside top regional talent, with team members from Talabat, Careem, Etisalat, and more