Job Overview
Job description
- Location:
- Brazil
- Work arrangement:
- Remote
Role Summary
100% remote vacancy
Here we connect the world
Sensedia is a Brazilian multinational provider of API management platforms, agent governance and integration solutions, as well as professional services. The company drives the digital transformation of large companies by enabling more agile, modern and scalable architectures, thus enabling its customers to offer digital products and experiences to their entire ecosystem.
Working here means belonging to a plural, relaxed and innovative culture. It is for those who have the courage to go further, think and act outside the box. We prefer to apologize rather than ask for permission and we are always willing to transform and reinvent ourselves.
Our people are incredible and you can be part of it all. We are committed to ensuring a welcoming and respectful work environment.
What is the mission of the Position?
Keep the platform that supports Open Solutions operations in production standing, observable and predictable. It is a multi-tenant environment, in the cloud (AWS, Kubernetes, Terraform), where unavailability is not an operational inconvenience — it is a failure to comply with regulatory obligations, with a reporting window and deadline defined by the regulator. The central challenge is to shift the operation from reactive to preventive: building observability that anticipates problems instead of finding them later, and eliminating repetitive manual work through automation. The person in this position will be responsible for Open solutions, working alongside cloud and software architects, as well as developers, with shared responsibility for what is live.
Responsibilities
What will your day-to-day activities be?
Observability and monitoring: Build and evolve the monitoring of production environments, with a focus on anticipating failure; Define and implement indicators that represent real user experience, not just resource health; Reduce alert noise and false positives; Maintain dashboards that support technical decisions.
Automation and toil reduction: Identify and eliminate repetitive manual work: provisioning, configuration, operational routines; Evolve infrastructure as code (Terraform, Helm) as a single source of truth; Develop scripts and automations in Python and Shell for recurring operations; Convert undocumented procedure into executable runbook.
Incident support and response: Act in the triage, mitigation and resolution of production incidents; Document and support post-mortems with a focus on root cause and preventive action; Support the environments with the engineering team; Monitor infrastructure consumption and costs, generating visibility for technical decisions
Here you will find
Meal Voucher/Food Voucher (Flash benefit card), Health Plan, Dental Plan, Life Insurance, PPR, TotalPass, Daycare Assistance, Well-Being Program (designed for physical and mental health), Corporate University (our # Sensedia Academy), with several development paths; Cultural and educational partners, with special discounts; We are a corporate citizen, providing maternity leave and extended paternity leave.
- We have #WorkWhereYouBelong as a value proposition, which is a flexible work model that helps us increase Sensediers' sense of belonging.
- Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know.
Requirements
Demonstrable experience, conceptual unfamiliarity
Have operated Kubernetes in production: pod debugging, resource consumption analysis, understanding failure modes Cloud in a mission-critical environment — AWS preferred
- Have written and maintained infrastructure as code with Terraform and Helm, including reviewing someone else's change
- Have built monitoring with Prometheus and Grafana: creating metrics and alert rules, not just consuming a ready-made dashboard
- Linux troubleshooting with autonomy: services, processes, network, disk, log analysis
- Automations in Python or Shell that other people have started using
- Participation in response to production incidents, with clarity on when to resolve and when to escalate
- Sufficient network knowledge to investigate: segmentation, DNS, TLS, VPN, debugging tools
What will be the different requirements for this position?
- Experience in a regulated environment: finance, insurance or similar
- Experience with APIs, REST and API gateway
- Log management at scale (Elasticsearch, OpenSearch, Loki or equivalent)
- Building and maintaining CI/CD pipelines
- Practice defining SLI and SLO, and using error budget to guide decisions
- Configuration as code (Ansible or equivalent) English for reading technical documentation
- Role:
- Cloud Operations Analyst (SRE) | Full (13443)
- Job Type:
- Full Time