United States
2 months ago

Job Overview

Job Type
Full Time
Pay
Not disclosed

Job description

Location:
United States
Work arrangement:
Hybrid

Role Summary

Join a well-funded, Series A AI startup building the next generation of autonomous site reliability engineering for the enterprise. Backed by top-tier investors and trusted by some of the largest companies in the world, this team is tackling one of the hardest problems in AI: autonomously detecting, diagnosing, and remediating complex production incidents in real time.

As an AI Engineer on the Data Platform team, you'll design, build, and maintain the backend systems that power an AI-driven observability platform. This hands-on role blends distributed systems engineering, low-level system design, performance optimization, observability, and AI integration — across both cloud and on-premises deployments.

Responsibilities

  • Architecture & Implementation: Contribute to the design and implementation of scalable, resilient infrastructure systems powering AI-driven root cause analysis and observability workflows, including on-premises deployment environments.
  • Low-Level System Design: Work on the foundational building blocks of the infrastructure, ensuring efficient resource utilization and high performance at scale.
  • Performance Optimization: Profile and tune backend systems to improve throughput, reduce latency, and eliminate bottlenecks across the stack.
  • Observability Systems: Build and maintain the internal observability stack — logs, metrics, and traces — used by AI agents to understand and act on production issues.
  • Hybrid Infrastructure: Support cloud and on-premises architecture to serve both SaaS and enterprise customer deployment models.
  • Cross-functional Collaboration: Work closely with engineers across the company to deliver resilient infrastructure that enables AI agents to diagnose and remediate production incidents in real time.

Requirements

  • Experience: 2–5 years of hands-on backend or infrastructure engineering experience.
  • Distributed Systems: Strong understanding of distributed systems design principles and trade-offs.
  • Performance Engineering: Proven experience profiling and optimizing high-throughput, low-latency systems.
  • Observability: Familiarity with observability tooling and concepts (logs, metrics, traces); experience with platforms such as Datadog, Grafana, Splunk, or similar is a plus.
  • Cloud & On-Prem: Experience with hybrid or multi-environment infrastructure (cloud + on-premises).
  • AI/ML Integration: Interest in or experience building systems that support AI/ML workloads at scale.
  • Background: Prior experience at observability, incident management, or data infrastructure companies is highly valued.
  • Note: Visa sponsorship is not available for this role.
  • This is a fully on-site role based in New York, NY . Remote work is not available for this position.
Role:
AI Engineer - Data Platform
Job Type:
Full Time

Company profile

Clera, legally Clera Labs, Inc., is an AI talent agent that connects job seekers directly with hiring managers at high-growth startups funded by investors such as a16z, Index, YC and GC. Instead of relying on mass applications, Clera learns what a person has done and wants next, then introduces them to the person who needs exactly that, and it shows candidates where they rank against people with comparable experience. The company began in a hacker house in Medellín and was co-founded by Alexander Farr, Sebastian Scott and Daniel Wintermeyer.

Headquarters
San Francisco, California, United States
Founded
2025
Founders
Sebastian Scott, Alexander Farr, Daniel Wintermeyer

More jobs at Clera

Similar Machine Learning Engineer jobs at other companies