HungryJob 🇵🇭
BC

Site Reliability Engineer (SRE)

Blackfort Consulting, Inc..

Santa Ana, Metro ManilaPosted 1 hour agoJobStreet

Setup
On-site
Type
Full-time
Level
Mid-level
Salary
Not listed
Closes
Open

Skills mentioned

PythonSQLAWSAzureDockerKubernetesCI/CDLinuxGitTerraformProduct ManagementLegalLeadership
Sign up free to tailor my resume for this job Apply on JobStreet

Members see a match score against their own skills and get a resume + cover letter written for this posting.

About the role

Reliability Engineering & Operations

Drive the adoption of Site Reliability Engineering (SRE) principles, including Service Level Objectives (SLOs), Service Level Indicators (SLIs), Error Budgets, monitoring, and alerting.

Design and implement automation solutions that improve operational efficiency, reduce manual effort, and increase service reliability.

Perform capacity planning, performance analysis, and system design to ensure platforms can scale effectively with business growth.

Lead troubleshooting efforts for complex production issues and perform root cause analyses (RCA) to prevent recurrence.

Participate in incident management processes, including major incidents, service outages, and on-call support when required.

Conduct post-incident reviews and drive corrective and preventive actions to continuously improve platform stability.

Leverage operational metrics and data-driven insights to proactively identify and mitigate potential risks.

Support engineering teams in adopting tools, processes, and best practices that improve reliability, maintainability, scalability, and extensibility.

Collaborate across Enterprise Platform teams to ensure standardized platforms meet enterprise-grade reliability and performance standards.

Champion the implementation of Non-Functional Requirements (NFRs), including availability, performance, security, and resilience, throughout the software development lifecycle.

Lead deep-dive investigations into problematic applications and services, identifying opportunities for architectural and operational improvements.

Oversee vulnerability management, platform lifecycle governance, and End-of-Life (EOL) remediation activities across supported products and services.

Leadership & Strategic Responsibilities

As a senior member of the team, you will

Lead and coordinate strategic initiatives and key deliverables within the Platform SRE function.

Serve as an advocate for SRE principles and operational excellence across the Enterprise Platform organization and broader business community.

Review engineering practices, operational processes, and team outputs to drive continuous improvement.

Foster strong collaboration with development, infrastructure, security, and business teams.

Contribute to the long-term strategy, roadmap, and maturity of the Platform SRE function.

Required Qualifications

Bachelor's degree in Computer Science, Information Technology, Software Engineering, or a related discipline.

Practical experience defining, implementing, and managing SLOs, SLIs, and Error Budgets.

Minimum of 5 years of experience supporting and operating production environments in SRE, DevOps, Platform Engineering, or related roles.

Experience or familiarity with monitoring and observability tools (e.g., Prometheus, preferably Grafana)

Professional fluency in English.

Proven ability to collaborate effectively within cross-functional teams.

Strong commitment to continuous learning, operational excellence, and continuous improvement.

Preferred Qualifications

Experience with Infrastructure as Code (IaC) and GitOps practices.

Knowledge of CI/CD platforms such as GitHub Actions, Jenkins, Azure DevOps, or GitLab.

Experience implementing Auto-Remediation and AIOps capabilities.

Understanding of security best practices, vulnerability management, and compliance frameworks.

AWS, Kubernetes, Terraform, or SRE-related certifications are advantageous.

Hands-on experience with observability, monitoring, logging, and alerting platforms such as Dynatrace, Datadog, Splunk, or similar tools.

Experience building automation solutions using technologies such as Python, Shell Scripting, PowerShell, Terraform, Ansible, Chef, Puppet, SQL, or equivalent.

Experience supporting middleware and platform technologies, including databases, web servers, messaging systems (MQ), and Kafka.

Familiarity with containerization and orchestration technologies such as Docker and Kubernetes.

Sourced from JobStreet · posted 1 hour ago · you apply on the original site