HungryJob 🇵🇭
BC

Platform Site Reliability Engineer (SRE)

Blackfort Consulting, Inc..

Santa Ana, Metro ManilaPosted 2 days agoJobStreet

Setup
On-site
Type
Full-time
Level
Mid-level
Salary
Not listed
Closes
Open

Skills mentioned

PythonSQLAWSAzureDockerKubernetesCI/CDLinuxGitTerraformProject ManagementProduct ManagementLegalNetworkingLeadership
Sign up free to tailor my resume for this job Apply on JobStreet

Members see a match score against their own skills and get a resume + cover letter written for this posting.

About the role

Reliability Engineering & Operations

Drive the adoption of Site Reliability Engineering (SRE) principles, including Service Level Objectives (SLOs), Service Level Indicators (SLIs), Error Budgets, monitoring, and alerting.

Design and implement automation solutions that improve operational efficiency, reduce manual effort, and increase service reliability.

Perform capacity planning, performance analysis, and system design to ensure platforms can scale effectively with business growth.

Lead troubleshooting efforts for complex production issues and perform root cause analyses (RCA) to prevent recurrence.

Participate in incident management processes, including major incidents, service outages, and on-call support when required.

Conduct post-incident reviews and drive corrective and preventive actions to continuously improve platform stability.

Leverage operational metrics and data-driven insights to proactively identify and mitigate potential risks.

Support engineering teams in adopting tools, processes, and best practices that improve reliability, maintainability, scalability, and extensibility.

Collaborate across Enterprise Platform teams to ensure standardized platforms meet enterprise-grade reliability and performance standards.

Champion the implementation of Non-Functional Requirements (NFRs), including availability, performance, security, and resilience, throughout the software development lifecycle.

Lead deep-dive investigations into problematic applications and services, identifying opportunities for architectural and operational improvements.

Oversee vulnerability management, platform lifecycle governance, and End-of-Life (EOL) remediation activities across supported products and services.

Leadership & Strategic Responsibilities

As a senior member of the team, you will

Lead and coordinate strategic initiatives and key deliverables within the Platform SRE function.

Provide technical leadership and influence architectural and design decisions across the organization.

Mentor, coach, and support the professional development of fellow engineers.

Identify emerging technologies, industry trends, and opportunities for innovation and continuous improvement.

Serve as an advocate for SRE principles and operational excellence across the Enterprise Platform organization and broader business community.

Review engineering practices, operational processes, and team outputs to drive continuous improvement.

Foster strong collaboration with development, infrastructure, security, and business teams.

Contribute to the long-term strategy, roadmap, and maturity of the Platform SRE function.

Required Qualifications

Bachelor's degree in Computer Science, Information Technology, Software Engineering, or a related discipline.

Minimum of 5 years of experience supporting and operating production environments in SRE, DevOps, Platform Engineering, or related roles.

Practical experience defining, implementing, and managing SLOs, SLIs, and Error Budgets.

Strong understanding of Linux and/or Windows system administration, networking fundamentals, and distributed systems concepts.

Hands-on experience with observability, monitoring, logging, and alerting platforms such as Datadog, Splunk, Grafana, Prometheus, or similar tools.

Experience building automation solutions using technologies such as Python, Shell Scripting, PowerShell, Terraform, Ansible, Chef, Puppet, SQL, or equivalent.

Working knowledge of cloud computing platforms, preferably AWS.

Experience supporting middleware and platform technologies, including databases, web servers, messaging systems (MQ), and Kafka.

Familiarity with containerization and orchestration technologies such as Docker and Kubernetes.

Excellent analytical, troubleshooting, and problem-solving skills.

Strong communication and stakeholder management abilities.

Ability to prioritize effectively and perform under pressure in fast-paced, mission-critical environments.

Professional fluency in English.

Proven ability to collaborate effectively within cross-functional and geographically distributed teams.

Strong commitment to continuous learning, operational excellence, and continuous improvement.

Preferred Qualifications

Experience with Infrastructure as Code (IaC) and GitOps practices.

Knowledge of CI/CD platforms such as GitHub Actions, Jenkins, Azure DevOps, or GitLab.

Experience implementing Auto-Remediation and AIOps capabilities.

Understanding of security best practices, vulnerability management, and compliance frameworks.

AWS, Kubernetes, Terraform, or SRE-related certifications are advantageous.

Sourced from JobStreet · posted 2 days ago · you apply on the original site