J
Team Leader, SRE at Remote
JobSearch
Nairobi 🇰🇪 Kenya
Posted 2 days ago (25 Sep 2026) Updated about 22 hours ago
Job Description
The Team Leader for Site Reliability Engineering will oversee the reliability of critical infrastructure, including Kubernetes, AWS, and CI/CD pipelines. They will guide a team of engineers, drive incident response processes, and align reliability practices with company objectives. The role blends hands‑on technical expertise with leadership responsibilities.
Qualifications & Requirements
- • Proven hands‑on experience in site reliability, DevOps, or cloud infrastructure engineering
- • Deep expertise with Kubernetes in production environments
- • Strong knowledge of AWS at scale
- • Experience building and scaling AI or machine‑learning infrastructure
- • Proficiency with infrastructure‑as‑code tools such as Terraform
- • Familiarity with CI/CD systems (GitLab CI, GitHub Actions, Jenkins) and container technologies (Docker, shell scripting)
- • Demonstrated ability to lead and develop engineering teams
- • Excellent communication, conflict‑resolution, and stakeholder‑management skills
- • Experience working in regulated environments is a plus
Required Skills
Responsibilities
- • Own and maintain Kubernetes, AWS, PostgreSQL, and CI infrastructure for the organization
- • Lead the SRE team, balancing individual contributor work with people‑management duties
- • Set technical direction, prioritize reliability initiatives, and align them with business goals
- • Represent the team in cross‑functional discussions and act as the spokesperson for reliability matters
- • Stay engaged with day‑to‑day engineering tasks to provide credible guidance and early issue detection
- • Implement and evolve SLO frameworks, error budgets, and incident‑response processes
- • Foster a culture of continuous improvement by turning incidents into lasting changes
- • Ensure observability, monitoring, and alerting practices meet production standards