Job Overview
Job description
- Location:
- Brazil
- Work arrangement:
- On-site
Role Summary
We are a company that uses technology to revolutionize healthcare, making the professional world more efficient and adding real value to the lives of thousands of people. Since 2017, we have been developing innovative solutions for doctors and healthcare professionals. Initially we simplified bureaucracy, now we follow the entire professional trajectory, from the beginning to retirement. We have more than 40 thousand healthcare professionals using our solutions every day. With a team of more than 400 employees throughout Brazil, our mission is clear: to facilitate the daily journey of healthcare professionals and transform the way we care for people 💙
Design, evolve and maintain multi-stage CI/CD pipelines with GitHub Actions, optimizing build time, cache and deployment strategies (blue/green, canary, rollback).
Develop and maintain reusable Terraform modules, with state management, workspaces and plan review via Atlantis (GitOps for IaC).
Operate and evolve environments in AWS (ECS/Fargate, EKS, Lambda, RDS/Aurora, S3, IAM, VPC and networking) and GCP (GKE, Cloud Run), focusing on availability, security and cost.
Deploy and sustain deliveries in Kubernetes with Helm and ArgoCD, including troubleshooting workloads, resources and cluster networking.
Implement and evolve the observability stack (Datadog, Prometheus, Grafana, Loki/Tempo, CloudWatch): dashboards, monitors, distributed tracing and log strategy.
Define and monitor SLIs/SLOs and error budgets for critical services, proposing concrete reliability actions based on the data.
Act as primary on-call scale: respond to incidents, conduct troubleshooting in production, write post-mortems without blame and address corrective actions.
Automate operational tasks and reduce toil through internal scripts and tools (Python, Go or Bash), including evolution of the team's internal CLI.
Integrate DevSecOps practices into pipelines (Trivy, Checkov, SonarQube, SAST/DAST), treating findings with development teams.
Contribute to FinOps initiatives: cost analysis and apportionment, tagging, right-sizing of resources and identification of savings opportunities.
Maintain and evolve the edge layer and network security (Cloudflare, WAF, DNS, ZTNA), in a context of sensitive health data subject to LGPD.
Mentor people at entry levels, participate in design/architecture reviews, define technical standards and document decisions, runbooks and architectures.
Use AI tools in the workflow to accelerate automation, troubleshooting, IaC review and documentation, with critical judgment on the generated output and care with sensitive data.
Collaborate directly with product and development teams, ensuring safe, automated, observable and scalable deliveries.
Mandatory
Solid experience in operating production environments on AWS, with autonomy in ECS/EKS, Lambda, IAM and networking (VPC, subnets, security groups, load balancers).
Mastery of Terraform beyond basic use: creation of modules, state management, workspaces and good provider versioning practices.
Hands-on experience building and maintaining complex pipelines in GitHub Actions.
Real experience with containers and orchestration (Docker, ECS and Kubernetes), including troubleshooting running applications.
Consistent knowledge in observability: metrics, logs and traces, creation of dashboards and actionable alerts (Datadog, Grafana, Prometheus or equivalent).
Linux administration and troubleshooting in production: processes, network, permissions, performance and resource analysis.
Scripting for automation (Python, Go or Bash) and mastery of Git and branching strategies.
Security fundamentals applied to infrastructure: secret management, principle of least privilege, hardening and network segmentation.
Active use of AI assistants in day-to-day technical work (e.g.: Claude Code, GitHub Copilot, Cursor) to write automations, investigate incidents and document, knowing how to critically validate the result and recognize the limits of the tool.
Ability to act autonomously on medium complexity projects, document what you build and communicate technical decisions clearly.
Requirements
Desirable
- AWS Solutions Architect Associate (SAA-C03) or GCP Associate Cloud Engineer certification — part of the team's career path and supported by the company.
- Additional certifications: Terraform Associate, CKA/CKAD, AWS DevOps Professional.
- Experience with GitOps (ArgoCD or Flux) and Helm in production.
- Experience with GCP (GKE, Cloud Run) and multi-cloud scenarios.
- SRE practices: definition of SLIs/SLOs, error budgets and conducting post-mortems.
- FinOps: cost analysis, tagging and optimization of cloud resources.
- Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know.
- Role:
- DevOps Engineer Pleno
- Job Type:
- Full Time