Brazil
1 month ago

Job Overview

Job Type
Full Time
Pay
Not disclosed

Job description

Location:
Brazil
Work arrangement:
On-site

Role Summary

  • We are Quality Digital! Find out more about us:
  • A phrase that defines us - We are experts in IT solutions and passionate about innovation! 🚀
  • To infinity and beyond - We are #withoutborders. Our team is spread across Brazil and the world 🌎

Our culture - Even though we are far apart, we are together 🧡 We have ceremonies to be closer, share knowledge and stay up to date with the latest news about our company!

We are diverse - One of our commitments is to make the company an increasingly diverse, inclusive and respectful environment, valuing and promoting plurality. Therefore, here we have space for everyone, regardless of race, gender, age, sexual orientation, religious belief or disability. 🏳️🌈👩🏾🦽

Our objective - Enhance our customers' businesses through innovative and sustainable solutions. Would you like to come with us on this one? 🎯

Experience with observability tools and platforms such as

  • Datadog
  • Elastic Stack / Elasticsearch / Logstash / Kibana / Elastic Agent
  • Sensedia / API Management
  • Zabbix
  • OpenTelemetry
  • Prometheus
  • Grafana
  • There will be differences

Intermediate level English.

  • Operating hours
  • 100% remote
  • 08:00 to 17:30 (lunch break from 11:30 to 13:00) - Temporary

NOTE: Hey! If you do not meet all the vacancy requirements, we invite you to apply anyway, ok?

We will carefully analyze your profile considering all your qualifications 😉

Responsibilities

  • Work with the Infrastructure, Networks, Database, Cloud, Security, Development, Architecture, Integration, Middleware, APIs and Systems teams, conducting technical surveys and discovery workshops.
  • Map the architecture and chain of dependencies of the company's main systems, identifying all components involved in the delivery of each service.
  • Build end-to-end Service Mapping, covering infrastructure, applications, integrations, APIs, databases, queues, external services and other dependencies.
  • Identify the different paths used by integrations between systems and understand how each component impacts service availability.
  • Relate technical components with the respective business services, allowing greater understanding of the impact of failures.
  • Identify existing monitoring and observability gaps in environments.
  • Define and implement an observability strategy for the mapped services, covering metrics, logs, traces, events, digital experience and availability.
  • Create service-oriented monitoring, moving from observing isolated components to monitoring the complete flow of transactions and integrations.
  • Define reliability indicators such as SLI, SLO and service availability.
  • Build dashboards, alerts, correlations and diagnostic mechanisms that allow you to quickly identify where degradation or interruption has occurred.
  • Work on reducing MTTD and MTTR, using observability to accelerate root cause identification.
  • Participate in corporate projects from their initial phases, ensuring that new systems and integrations are born with adequate observability requirements.
  • Support the evolution of the observability culture within the organization, promoting good practices among different technical teams.
  • Complete higher education required;
  • Application architecture and distributed systems
  • Server infrastructure and operating systems
  • Communication networks and protocols
  • Databases
  • REST APIs and cross-system integrations
  • Microservices
  • Containers and Kubernetes
  • Cloud and hybrid environments
  • Queues and messaging
  • Balancers, proxies and gateways
  • DNS, HTTP/HTTPS, TCP/IP and TLS
  • APM and Distributed Tracing
  • Log centralization and analysis
  • Infrastructure metrics and monitoring
  • Service Mapping and Dependency Mapping
  • SLI, SLO, SLA and Error Budget
  • Incident Management and root cause analysis

What you will find here

  • An environment conducive to learning and professional growth 🎯
  • Performance assessment and feedback, aiming for the continuous development of our people 📊
  • Food and/or meal voucher for your grocery shopping and meals 🍴
  • Medical and dental assistance to keep you and your family in good health 💙
  • Agreement with pharmacies for discounts on medicines 💊
  • Childcare assistance in accordance with current policy 🍼
  • Partnership with SESC for varied cultural and leisure programs ✈
  • Partnerships for your language studies, technology and course platform 📚
  • Payroll loan with attractive rates + financial education program💰
  • Corporate University and knowledge trails with diverse technology content, soft skills, market trends and much more 👨💻
  • Referral Program with the possibility of prizes and bonuses 🎁
  • Group life insurance ⛑

Requirements

  • We are looking for a Site Reliability Engineering professional with strong specialization in Observability, capable of working across different areas of technology to map, understand and monitor the organization's complete chain of services and systems.
  • The main challenge of this position will be to build an end-to-end view of services, identifying how applications, integrations, APIs, databases, infrastructure components, networks, messaging and third-party services relate to each other to deliver each business service.
  • Strong capacity for technical investigation and systemic vision, making it easy to talk to experts from different disciplines and transform fragmented information into an integrated view of the environment.
Role:
SRE SPECIALIST
Job Type:
Full Time

More jobs at Quality Digital

Similar DevOps Engineer jobs at other companies