Submitting more applications increases your chances of landing a job.

Here’s how busy the average job seeker was last month:

Opportunities viewed

Applications submitted

Keep exploring and applying to maximize your chances!

Looking for employers with a proven track record of hiring women?

Click here to explore opportunities now!
We Value Your Feedback

You are invited to participate in a survey designed to help researchers understand how best to match workers to the types of jobs they are searching for

Would You Be Likely to Participate?

If selected, we will contact you via email with further instructions and details about your participation.

You will receive a $7 payout for answering the survey.


https://bayt.page.link/Ly2AQMQFz1bbzRrF7
Back to the job results

Site Reliability Engineer

30+ days ago 2026/02/21
Other Business Support Services
Create a job alert for similar positions
Job alert turned off. You won’t receive updates for this search anymore.

Job description

Roles and Responsibilities

Roles and Responsibilities:
* Design and implement the lifecycle of services from conception to inception, including system design, build, and deployment
* Develop software solutions to enable operability of large-scale distributed systems capable of handling millions of transactions and petabytes of data
* Manage capacity and performance to help scale the infrastructure both on public and private clouds around the world
* Define and implement standards and best practices related to: System Architecture, Deployment, metrics, operational tasks
* Support services through activities such as monitoring availability, system health, and incident response
* Improve system performance, application delivery and efficiency through automation, process refinement, postmortem reviews, and in-depth configuration analysis
* Engage in Communications across all areas of the organization
* Troubleshooting and monitoring production systems to ensure the highest uptimes are maintained
* Support and improve upon existing high-availability architecture solutions as well as manage the operational activity.
* Integrate Generative AI (GenAI) and AIOps tools to automate incident detection, root cause analysis, and resolution workflows (e.g., self-healing scripts, intelligent runbooks), reducing manual toil and accelerating response times.
* Apply Prompt Engineering techniques to enhance interactions with AI-based observability and automation platforms improving accuracy and efficiency of AI responses.
* Leverage platform-specific AI capabilities (e.g., AWS Bedrock, Azure OpenAI, GCP Vertex AI) to architect intelligent SRE solutions tailored to cloud environments.
* Design, implement, and maintain AI/ML driven monitoring and alerting systems to proactively detect anomalies and predict potential failures, enabling preemptive remediation.
* Develop and train machine learning models using operational telemetry (logs, metrics, events, traces) to support predictive analytics and intelligent automation.
* Evaluate and deploy AIOps platforms (e.g., Moogsoft, Dynatrace, Splunk, BigPanda, Datadog, Elastic) to enhance observability, reduce noise, and accelerate incident resolution.
* Experience in one or more high level programming languages like Python or Ruby or GoLang and familiar with Object Oriented Programming.


This job post has been translated by AI and may contain minor differences or errors.

You’ve reached the maximum limit of 15 job alerts. To create a new alert, please delete an existing one first.
Job alert created for this search. You’ll receive updates when new jobs match.
Are you sure you want to unapply?

You'll no longer be considered for this role and your application will be removed from the employer's inbox.