Philippines
1 month ago

Job Overview

Job Type
Full Time
Pay
Not disclosed

Job description

Location:
Philippines
Work arrangement:
On-site

Role Summary

Our client is a technical consulting company specializing in operational services for the high-tech industry. They specialize in helping platform and infrastructure teams operate multi-cloud environments, execute complex migrations, and enable seamless app deployments.

Incident & Problem Management

Lead end-to-end incident response, triage, communication, and resolution in real time.

Act as Service Operations & Reliability Incident Commander for high-impact events across a global environment.

Track and improve metrics like MTTD, MTTM, and MTTR.

Champion blameless Post-Incident Reviews (PIRs) and translate learnings into long-term system and process improvements.

Service Operations & Reliability

Oversee daily service health, capacity, and reliability across all supported environments.

Ensure compliance with operational KPIs through proactive planning and improvement.

Balance demand vs. capacity and manage shift coverage to prevent burnout.

Partner with engineering teams to maintain runbooks, knowledge bases, and escalation paths.

Drive automation and workflow optimization to reduce manual overhead.

Use data insights to guide decisions and improvements.

Strategic & Cross-Functional Impact

Represent in customer reviews, operational syncs, and briefings.

Collaborate with SREs, product owners, and partner engineers to align priorities and reliability goals.

Contribute to frameworks and governance initiatives.

Lead service onboarding/off-boarding and strengthen operational readiness checkpoints.

Identify and close systemic operational gaps through process and tool improvements.

Requirements

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline.
  • 3+ years in Service Delivery, Incident Response, or Operations Leadership within enterprise-scale, 24×7 environments.
  • Proven experience managing technical teams, driving performance, and leading through critical situations.
  • Strong grounding in ITSM / ITIL principles (Incident & Problem Management).
  • Familiarity with cloud, distributed systems, or enterprise infrastructure.
  • Skilled in monitoring, alerting, and ticketing tools (e.g., PagerDuty, Datadog, Grafana, Splunk, ServiceNow).

Core Competencies

  • Incident Command and Escalation Management
  • Analytical and Problem-Solving Skills
  • Communication and Decision-Making Under Pressure
  • Root Cause and Post-Incident Analysis
  • Operational Planning and Service Governance
  • Stakeholder and Partner Management
  • IT Service Management (Incident & Problem Management)
  • Observability, Monitoring, and Automation Tools

Plus points if you have

ITIL V3 or V4 certification or Incident Management Certifications

Familiarity in SRE practices and operational frameworks that promote reliability and automation

Role:
Incident Response Lead
Job Type:
Full Time

Company profile

Permhunt

permhunt.com

Permhunt is a technology recruitment agency that helps employers around the world hire software engineers and other technical staff from the Philippines. It sends first CVs within two business days, charges one flat fee with no markup on salaries, and replaces a hire free of charge if it does not work out. Permhunt pairs recruitment expertise with an in-house engineer and keeps a database of developer, cybersecurity, and DevOps candidates for startups, agencies, and enterprises.

More jobs at Permhunt

Similar Security Engineer jobs at other companies