1. Home
  2. Jobs
  3. Acesoft Labs (India) Pvt Ltd
  4. Site Reliability Engineer
AI

Site Reliability Engineer

Acesoft Labs (India) Pvt Ltd
Dubai, UAE Listed 1h ago via Naukrigulf
python docker kubernetes aws azure gcp terraform ansible jenkins github actions ci/cd devops linux helm machine learning

Job Description Roles & Responsibilities SRE / AIOps Engineer Experience 4–8 Years Location UAE – Dubai / Abu Dhabi Employment Type Full-Time / Contract Work Mode Onsite / Hybrid – As per client requirement Role Overview We are looking for an experienced SRE / AIOps Engineer to ensure the reliability, availability, scalability, and performance of enterprise applications and cloud infrastructure. The ideal candidate should have strong hands-on experience in Site Reliability Engineering, cloud platforms, Kubernetes, observability, monitoring, automation, incident management, and CI/CD. Exposure to AIOps, machine learning-based anomaly detection, predictive monitoring, and automated incident remediation will be highly preferred. Key Responsibilities Design and implement highly available, scalable, and resilient infrastructure and applications. Establish and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs). Monitor application and infrastructure health, performance, availability, and capacity. Implement comprehensive observability solutions across applications, infrastructure, networks, and cloud environments. Develop dashboards, alerts, metrics, logs, and distributed tracing solutions. Identify performance bottlenecks and conduct root-cause analysis. Automate incident detection, diagnosis, and remediation. Implement AIOps capabilities to identify anomalies, correlate events, and reduce operational noise. Develop automated workflows for recurring operational and infrastructure tasks. Support production incidents and participate in on-call / 24x7 support when required. Conduct post-incident reviews and implement corrective and preventive actions. Improve system reliability and reduce Mean Time to Detect (MTTD) and Mean Time to Recovery (MTTR). Implement capacity planning, performance engineering, and disaster-recovery strategies. Work closely with Development, DevOps, Cloud, Security, and Infrastructure teams. Implement Infrastructure as Code and configuration automation. Build and maintain CI/CD pipelines for reliable and repeatable deployments. Continuously identify opportunities to improve system reliability through automation and AI. Mandatory Skills 4–8 years of experience in SRE / DevOps / Cloud Engineering. Strong experience with Linux/Unix systems. Strong experience with AWS, Azure, or GCP. Hands-on experience with Kubernetes and Docker. Strong experience with monitoring and observability tools. Experience with Prometheus and Grafana or equivalent. Experience with centralized logging platforms such as ELK / OpenSearch. Experience with CI/CD tools such as Jenkins, GitLab CI, GitHub Actions, or Azure DevOps. Strong scripting/programming skills in Python, Bash, or PowerShell. Experience with Terraform / Infrastructure as Code. Strong understanding of networking, cloud infrastructure, and distributed systems. Experience with incident management and root-cause analysis. Good understanding of SLI, SLO, SLA, availability, latency, scalability, and reliability concepts. Observability Experience with one or more: Prometheus Grafana OpenTelemetry Splunk ELK / Elastic Stack OpenSearch Datadog Dynatrace New Relic AppDynamics Knowledge of metrics, logs, traces, distributed tracing, and application performance monitoring (APM) is preferred. AIOps Capabilities Candidates with hands-on experience in the following will be preferred: AI-driven event correlation Automated anomaly detection Predictive monitoring Intelligent alerting Noise reduction Automated incident classification Root-cause analysis using AI Predictive capacity planning Automated remediation AI-assisted troubleshooting IT operations automation Cloud Technologies AWS EC2 EKS CloudWatch Lambda S3 IAM Auto Scaling CloudFormation Azure Azure Monitor AKS Application Insights Azure Automation Log Analytics Azure DevOps GCP GKE Cloud Monitoring Cloud Logging Cloud Trace Cloud Operations Suite Kubernetes & Containerization Strong understanding of: Kubernetes architecture EKS / AKS / GKE Docker Helm Kubernetes monitoring Resource management Auto-scaling Health checks Service discovery Kubernetes networking Container troubleshooting Infrastructure as Code & Automation Experience with: Terraform Ansible CloudFormation / ARM / Bicep Python Bash / PowerShell Git GitOps Argo CD Incident & Reliability Management Experience with: Incident response Problem management Root Cause Analysis (RCA) Disaster Recovery Business Continuity Capacity planning Performance engineering Change management Production support On-call operations Good-to-Have Skills AIOps platforms such as Dynatrace Davis, Moogsoft, BigPanda, Splunk ITSI, or similar. OpenTelemetry. ServiceNow ITOM / ITSM. Kafka and distributed messaging systems. Microservices architecture. Service mesh technologies such as Istio. Chaos Engineering. FinOps / cloud cost optimization. DevSecOps. AI/ML fundamentals. Experience integrating LLMs into IT operations. Knowledge of MLOps / LLMOps. Experience with large-scale enterprise environments. Certifications – Preferred AWS Certified DevOps Engineer Microsoft Azure DevOps Engineer Google Professional Cloud DevOps Engineer Certified Kubernetes Administrator (CKA) Certified Kubernetes Application Developer (CKAD) Certified Kubernetes Security Specialist (CKS) HashiCorp Terraform certification ITIL certification SRE / DevOps certifications Education Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related discipline Desired Candidate Profile Candidate Profile The ideal candidate should: Have strong troubleshooting and analytical capabilities. Be comfortable working with complex distributed systems. Have a strong automation mindset. Understand reliability and availability engineering principles. Be able to work effectively during production incidents. Have strong communication and stakeholder-management skills. Be comfortable working in a fast-paced enterprise environment. Demonstrate an interest in using AI and automation to improve IT operations. Employment Type Full-time Company Industry RecruitmentPlacement FirmExecutive Search Department / Functional Area Software DevelopmentApplication Development (IT Software) Keywords SRESite Reliability Analyst Get real-time job updates only on our App

Ready to apply?

You are viewing this role on JobSphere AI. Applications are completed on the original employer / source website.

Apply Now

Opens the employer's site in a new tab

  • CompanyAcesoft Labs (India) Pvt Ltd
  • LocationDubai, UAE
  • CategoryAI
  • SourceNaukrigulf
  • Listed1h ago

Related AI jobs

More AI