Site Reliability Engineer - (CONTRACT)

 

Recruiter:

PM Connection

Job Ref:

2383262979

Date posted:

Wednesday, September 29, 2021

Location:

Johannesburg, South Africa

Salary:

Negotiable


SUMMARY:
-

POSITION INFO:

Key Purpose:

The Site Reliability Engineer is responsible for driving initiatives proactively leading to high service and platform availability, improved performance and customer experience, enhancing and optimizing monitoring coverage and working with cross-functional teams to proactively build and maintain more reliable services and platforms.

Â

Areas of responsibility may include but not limited to:

  • Design and implement an observability framework across infrastructure, application and services deployed that can be centrally configuration managed and be deployed to all environments.
  • Instrumenting specific java methods or querying values stored in Java Objects for test validation or specific Business metrics.
  • Work with cross-functional teams to identify, evaluate and establish initiatives for improvement to services or processes with the purpose of increased availability, improved service levels, reduced costs, and improved customer satisfaction by reducing the number of operational problems.
  • Participate in various cross-functional forums and lead work streams to contribute to the improvement and implementation of policies, frameworks and standards.
  • Responsible for driving initiatives regarding software automation and reliability.
  • Develop and optimize monitoring framework based on industry best practices by developing metrics (SLI`s), monitoring, and alerting (SLO`s) to observe the health of the production system.
  • Maximize value of tooling and leverage metrics for data driven insights and problem management.
  • Gather and analyse metrics from various monitoring tools covering but not limited to operating systems, infrastructure and applications to assist in performance tuning and fault finding.
  • Facilitate discussions (including technical discussions) to establish root causes and solutions to any infrastructure, application or process related issues.
  • Conduct research to establish more efficient ways of performing day to day activities using new technologies or frameworks and identify opportunities for automation.
  • Conduct trend analysis of data, both systematically and manually to determine common occurrences and recurring issues to feed into the Problem Management processes.
  • Perform impact assessments to determine priority of a problem relative to other problems and business activities.
  • Clear, concise and timely communication with emphasis on expressing technical issues in a non-technical manner to clients and executives.
  • Drive infrastructure and application performance and availability initiatives.
  • Produce and present regular reports on availability, capacity and service performance.
  • Participation and facilitation of Incident Post-mortems.
  • Build and integrate metric collectors such as Prometheus for required metrics.
  • Drive the improvement of availability and performance using quality gates of the build pipeline
  • Liaise with application development teams to improve services and customer experience.

Â

Personal Attributes and Skills:

  • Statistical analysis and reporting
  • Problem solving
  • Root Cause Analysis
  • Business writing (reports) and presentation
  • Tenacity
  • Stress Management
  • Persuasion
  • Coaching
  • Client orientation

Â

Education:

  • Relevant Tertiary qualification (Bachelors’ Degree in IT or Engineering)

Knowledge and Experience:

  • 5 or more years’ experience in a Software Engineering, DevOps Engineer, SRE or Architecture role
  • APM and Infrastructure Monitoring Tool Experience (Prometheus, DynaTrace and Cloudwatch beneficial)
  • Knowledge of Architecture Frameworks, Tools and Standards
  • Experience in Application Performance Monitoring, JVM profiling and Prometheus
  • Extensive experience managing complex and high-volume applications
  • Experience optimizing database, infrastructure and application configurations
  • Experience supporting microservices based applications on a Kubernetes platform
  • Experience with AWS technologies and event driven architecture
  • Experience with event driven orchestration

Â

Kindly regard your application as unsuccessful if you have not heard from the agency within 2 weeks.



 

NB! This job is now closed. You can apply for other jobs by uploading your CV.



 

 

 

Similar jobs you might be interested in:

Sales Engineer
Location: Pretoria
Salary: 400 000
Where engineering meets opportunity — sell solutions that keep mines moving.
2 days ago


Reliability Engineer
Location: Pretoria
Salary:
A leading mining and processing operation is seeking a self-driven reliability engineer to join its team and provide technical leadership in improving equipment reliability. This is a full-time, on-site role focused on implementing systematic planning, analysis, and reporting processes to enhance asset performance and reduce downtime.
3 days ago


Mechanical Design Engineer
Location: Highveld
Salary:
Company based in Highveld Mpumalanga is looking for a skilled Mechanical Design engineer to join their team.
44 days ago


Mechanical Design Engineer
Location: Ngwedi
Salary:
Company based in Ngwedi North West, is looking for a skilled Mechanical Design engineer to join their team.
44 days ago


DevOps Engineer
Location: Johannesburg
Salary: Market-related
Own resilient infrastructure at the heart of digital security!
58 days ago


Senior DevOps Engineer (remote)
Location: Johannesburg
Salary: Market related
Senior DevOps engineer (remote)
80 days ago


Operations Director
Location: Johannesburg
Salary:
Today


Engineering Manager
Location: East Rand, Gauteng
Salary: Salary, bonus
My client is seeking an energetic and robust engineering Manager (preferably with FMCG experience). A degree in engineering is required. Previous Millwright qualification is advantageous. Will act in the capacity of GMR2.1- relevant qualification is applicable. Min 10 years'''' experience.
Today


Automation Engineer
Location: Johannesburg
Salary: 450000 Annually
We’re recruiting on behalf of a specialist engineering and technology solutions provider seeking a remote condition monitoring specialist to support plant reliability, predictive maintenance, and operational uptime across an installed base of process plants and equipment. This role is critical in translating remote monitoring data into actionable technical and operational outcomes.
1 day ago


Microsoft .Net C# web developer
Location: Pretoria
Salary: Hourly
🌟 We're hiring Mid–Senior .NET C# Web Developer (Support & Maintenance)Location: Pretoria (On-site/ Office based)contract Duration: 12 months (01 Feb 2026 – 31 Jan 2027)Working Hours: Full-time (8 hours per day)About the RoleOur client in the banking industry, is seeking a Mid–Senior Microsoft .NET C# Web Developer to provide application support, maintenance, and enhanceme...
4 days ago


Create a free job alert for Site Reliability Engineer - (CONTRACT) in Johannesburg

Enter your email address below and we will email you similar jobs when they become available:

You can cancel at any time. We will not spam you.
By giving us your email address your agree to our Terms and Conditions