San Francisco
3 months ago

Job Overview

Pay
$150,000 - 200,000

Job description

Location:
San Francisco
Work arrangement:
On-site

Role Summary

We’re looking for a detail-oriented Data Center Operations professional to manage and track all break/fix activities across multiple data center locations. This role acts as the central point of coordination for hardware incidents, vendor dispatches, ticket management, asset tracking, and operational reporting to ensure maximum uptime and fast issue resolution.

Responsibilities

  • Track and manage all break/fix incidents across multiple data centers
  • Monitor ticket queues and ensure SLA compliance for incident response and resolution
  • Coordinate with on-site technicians, remote hands teams, vendors, and engineering groups
  • Maintain accurate records of failed hardware, replacements, RMAs, and repair status
  • Escalate critical outages and recurring infrastructure issues to leadership and engineering teams
  • Schedule and oversee maintenance windows and emergency repair activities
  • Provide daily/weekly operational status reports and incident summaries
  • Ensure all work follows data center operational procedures and change management policies
  • Identify trends in hardware failures and recommend process improvements

Requirements

  • Experience working in data center operations, IT infrastructure, or hardware support
  • Strong understanding of server, storage, and networking hardware
  • Experience with ticketing systems such as ServiceNow, Jira, or Remedy
  • Ability to manage multiple priorities across several sites simultaneously
  • Excellent communication and organizational skills
  • Familiarity with SLA management and incident escalation processes
  • Proficiency with Excel, reporting dashboards, and inventory tracking tools

Preferred Qualifications

  • Experience supporting enterprise or hyperscale data centers
  • Knowledge of remote hands operations and vendor management
  • Understanding of ITIL processes and change management
  • CompTIA Server+, Network+, or similar certifications

About Together AI

Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month.

Salary

We offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $150,000-200,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.

Additional Information

Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Please see our privacy policy at https://www.together.ai/privacy

Role:
Data Center Operations Coordinator

Company profile

Together AI

together.ai

Together AI operates an AI cloud platform for production AI, covering inference, fine-tuning, model shaping, pre-training and GPU clusters. Developers can run open-source models on demand without managing infrastructure or signing long-term commitments, and the company also offers the Together Kernel Collection. Together AI raised an $800 million Series C led by Aramco Ventures, and it rents out Nvidia GPU clusters and other AI-specific infrastructure.

Headquarters
San Francisco, California, United States
Founded
2022
Founders
Vipul Ved Prakash
Funding Stage
Series C

More jobs at Together AI

Similar Operations Manager jobs at other companies