Site Reliability Engineer - Azure, Observability and Scripting
Job description
Jalasoft is seeking a Site Reliability Engineer (SRE) to join our team and help ensure the reliability, scalability, and performance of cloud-native platforms running on Microsoft Azure and Kubernetes. The ideal candidate is passionate about observability, automation, and operational excellence, with experience improving system availability, monitoring, incident response, and infrastructure automation in production environments.
Required Years of Experience:
- 6+ years of experience
- 3+ years of experience operating Kubernetes in production
Must Haves:
- Site reliability engineering or production operations for Kubernetes workloads at scale.
- Azure Monitor, Log Analytics and KQL, including workspace design, data collection rules and retention strategy.
- Prometheus and Grafana: metrics and exporters, recording and alerting rules, and dashboard design.
- Definition and implementation of service level indicators, objectives and error budgets.
- Alerting and incident response design, including runbook authoring and on-call practice.
- Backup, restore and disaster recovery design and testing, including validation against RPO and RTO targets.
- Kubernetes operations: workload troubleshooting, resource management and cluster upgrades.
- Ability to read and modify infrastructure as code (Terraform or Bicep) and Azure DevOps pipelines.
- Scripting in Python, PowerShell or Bash.
- Professional working English.
Nice-to-Have:
- Azure Managed Prometheus and Azure Managed Grafana.
- OpenTelemetry instrumentation and distributed tracing.
- Azure Backup, Azure Site Recovery, and snapshot-based recovery of virtual machines.
- Chaos engineering or structured game day practice.
- Database-layer observability, particularly for Oracle.
- Cost and capacity management for AKS estates.
- Incident management tooling and postmortem practice.
- Certification: CKA, AZ-400 or equivalent.