Bakı, PBT2 (Port Baku Tower 2),
Texnologiya
Razılaşma ilə
27 iyul 2026
27 avqust 2026
The Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of critical systems by combining software engineering with infrastructure and operations expertise. The role focuses on automation, cloud infrastructure, monitoring, incident management, and close collaboration with engineering teams to deliver highly available and resilient services.
+ ' ' +Strong knowledge of computer science fundamentals (data structures, algorithms, OS, networking);
Software development experience with at least one modern language (e.g., Go, Python, Java, C#);
Experience designing and supporting distributed systems and microservices;
Proficiency in troubleshooting complex issues in production environments;
Hands-on experience with observability tools (e.g., Prometheus, Grafana, OpenTelemetry);
Working knowledge of Kubernetes and containerized infrastructure;
Familiarity with cloud platforms (AWS, GCP, Azure);
Solid understanding of Linux systems and scripting;
Experience with CI/CD pipelines and DevOps practices;
Strong problem-solving skills and a curious, analytical mindset;
Effective communicator - able to clearly explain technical concepts to both engineers and non-engineers;
Team player with a collaborative approach to working across engineering, product, and operations;
Takes ownership and initiative;
Comfortable working in a fast-paced, evolving environment;
Attention to detail while keeping an eye on the bigger picture;
Eager to learn continuously and stay up to date with emerging technologies and practices.
+ ' ' +Collaborate with development and infrastructure teams to design and maintain scalable, reliable systems;
Write automation and tooling for infrastructure management, deployments, and operational tasks;
Build and maintain observability stacks (metrics, logs, traces) to ensure visibility and fast issue resolution;
Lead and participate in troubleshooting and debugging efforts across the full stack - application, platform, and infrastructure;
Conduct post-incident analysis, drive root cause investigations, and implement long-term fixes;
Define and monitor service level objectives (SLOs), indicators (SLIs), and error budgets;
Contribute to CI/CD pipelines and infrastructure as code efforts. - Continuously seek to improve system performance, resilience, and developer experience
Kapital Bank iş mühiti, əlavə fürsətlər və digər vakansiyaları görüntüləmək üçün Kapital Bank Life səhifəsinə keçid edin.
Vakansiyalardan daha tez xəbərdar olmaq üçün Telegram kanalımıza abunə olun!
Site Reliability Engineer (SRE) vakansiyaları olduğu zaman anında bildirişi e-poçtunuza alın.
Sizin elan saytın ana səhifəsində xüsusi ayrılmış blokda görünəcək və xidmətin
aktivlik
müddətinin sonunadək orada qalacaq.
Bu əməliyyatı etmək üçün profilə giriş etməyiniz tələb olunur.
Bu əməliyyatı etmək üçün profilə giriş etməyiniz tələb olunur.