Site Reliability Engineer-Malaysia

Talent Potential Consulting

Date: 1 week ago

City: Kuala Lumpur

Contract type: Full time

About Client:

With a global network and a reputation for excellence, this leading firm in Malaysia delivers comprehensive professional services, including consulting, tax, audit, and advisory. Known for its commitment to innovation and fostering client growth, it serves a wide range of industries, offering tailored solutions that meet both local and international needs. Recognized for its expertise, the firm empowers clients to navigate complex business landscapes, supporting sustainable growth and long-term success.

Role Overview:

As a Site Reliability Engineer (SRE), you will play a key role in maintaining the reliability and performance of critical services. Your expertise will help bridge the gap between development and operations, ensuring robust, scalable, and responsive infrastructure. This role emphasizes strong system architecture and design principles, focusing on key SRE practices such as Service Level Objectives (SLOs), Service Level Indicators (SLIs), and the reduction of operational toil. You will collaborate closely with diverse teams to drive reliability improvements and foster a culture of continuous learning and accountability

PROFICIENCY IN MANDARIN LANGUAGE IS MANDATORY

Key Responsibilities:

Design and implement resilient system architectures that support high availability and scalability.
Develop automation tools and scripts to enhance operational efficiency and reduce manual effort.
Define, track, and analyze SLOs and SLIs to ensure reliability and performance meet business needs.
Conduct thorough post-mortem analyses following incidents, driving continuous improvement through root cause identification and solution implementation.
Collaborate with development and operations teams to establish best practices in system reliability and incident management.
Troubleshoot and resolve issues related to database performance, network connectivity, and deployment failures, including diagnosing problems at the underlying platform level (e.g., Kubernetes, virtual machines).
Ensure that issues are resolved within the stipulated Service Level Agreements (SLAs), maintaining high standards of service delivery.
Identify and troubleshoot performance bottlenecks across systems, providing actionable recommendations for enhancements.
Maintain detailed documentation of processes and incident responses to support knowledge sharing and compliance.

Qualifications

Minimum 6 months experience
proficiency in Mandarin Language is MANDATORY
Proficiency in programming languages such as Python, Golang, Java, or similar, focusing on operational efficiency.
Demonstrated experience in system architecture and design, prioritizing reliability, and scalability.
Strong understanding of SRE principles, including SLOs, SLIs, toil reduction, and incident post-mortems.
Experience with cloud environments (e.g., AWS, Azure, Google Cloud) and their operational management.
Strong expertise in Linux system administration.
Proven experience in troubleshooting application support issues with a focus on performance and connectivity.
Familiarity with networking concepts and effective troubleshooting techniques.
Excellent problem-solving abilities and a proactive approach to operational challenges.
Ability to work independently while effectively collaborating within a team environment.

Preferred Skills:

Familiarity with monitoring tools and performance optimization techniques.
Experience in scripting or automation for system administration tasks.
Knowledge of networking concepts and troubleshooting methodologies.
Hands-on knowledge of cloud platforms (e.g., AWS, Azure, Google Cloud) and their services.
Familiarity with DevOps practices and frameworks, including CI/CD, infrastructure as code, and containerization.

How to apply

To apply for this job you need to authorize on our website. If you don't have an account yet, please register.

Post a resume

Similar jobs

Digital Marketing and Sales Executive

Xair Energy Sdn Bhd, Kuala Lumpur

1 day ago

Job Responsibilities:Assist in the formulation of strategies to build a lasting digital connection with consumersPlan and monitor the ongoing company presence on social media (Website, Facebook etc.)Drive website traffic: Using tactics like SEO, PPC, and SEM, to boost website visibility, increase company and brand awareness.Manage the organization's website: This includes updating content, ensuring the design is up-to-date, and optimizing the...

Executive, Group Digitalisation (Network & IoT)

Ingress Industrial (Malaysia) Sdn Bhd, Kuala Lumpur

3 days ago

The OpportunityIngress Industrial (Malaysia) Sdn Bhd, a leading industrial conglomerate, is seeking an exceptional Executive to join our Group Digitalisation team. As the Executive, Group Digitalisation (Network & IoT), you will play a pivotal role in driving the digital transformation of our organization, with a focus on network infrastructure and Internet of Things (IoT) solutions.Key ResponsibilitiesSupport the SAP technical activities...

A4 Chargeman (Shift Work) - Kuala Lumpur

CBRE, Kuala Lumpur

5 days ago

A4 Chargeman (Shift Work) - Kuala Lumpur Job ID 192851 Posted 12-Nov-2024 Service line GWS Segment Role type Full-time Areas of Interest Building Management, Engineering/Maintenance, Facilities Management Location(s) Kuala Lumpur - Wilayah Persekutuan Kuala Lumpur - Malaysia, Petaling Jaya - Selangor - Malaysia About the Role: As a CBRE Chargeman, you will inspect, repair, and maintain mechanical and electrical equipment...