Brazil
1 month ago

Job Overview

Job Type
Full Time
Pay
Not disclosed

Job description

Location:
Brazil
Work arrangement:
On-site

Role Summary

Caju is a Brazilian technology company that seeks to add flavor to professional life, transforming the relationship between companies and employees through more innovative and secure solutions such as the Multi Benefits Card, Corporate Expense Solution, Awards and Caju Cycles.

Here at Caju, we always learn, and become better and better in a collaborative and fun environment!

Applications from black people, women, indigenous people, LGBTQIA+, or other minority groups are very welcome.

Sign up and learn more about our team 🧡

Build and maintain robust data modeling using DBT, applying good development practices such as testing, documentation and model versioning.

Develop fact tables, dimensions and data marts oriented to the needs of the business areas, following dimensional modeling standards (Star Schema/Snowflake Schema).

Act as a link between business teams and data engineering, carrying out requirements gathering, interpreting business rules and translating analytical demands into reliable and scalable data solutions.

Document the models developed, including metrics definitions, applied business rules, data glossary and lineage, ensuring traceability and understanding throughout the organization.

Implement and maintain orchestration pipelines in Databricks and Airflow, ensuring reliability and monitoring of data flows.

Develop solutions in Python and SQL for transforming, processing and aggregating data in the Databricks environment.

Implement data quality tests in DBT models and analytical layers, ensuring consistency, completeness and accuracy of information delivered to business areas.

Structure monitoring and alerts on pipelines and data models, acting proactively to identify and resolve inconsistencies.

Lead complex end-to-end modeling projects, from understanding business pain to delivering analytical layers into production.

Apply appropriate partitioning, clustering and materialization strategies to ensure performance and cost efficiency in data processing.

Collaborate with product and business squads to ensure that data solutions support the company's strategic and operational decisions.

Use GitHub for code versioning, reviewing pull requests and maintaining good engineering practices within the team.

Data Modeling and Architecture

Mastery in analytical data modeling: Star Schema, Snowflake and OBT (One Big Table)

Solid knowledge of layered architectures such as Medallion Architecture (Bronze, Silver, Gold).

Ability to define and enforce cross-tier data contracts, ensuring consistency and predictability for consumers.

Performance and Optimization

Experience with materialization strategies in DBT (tables, views, incremental models and snapshots), knowing how to choose the most appropriate approach for each layer and volume of data.

Knowledge of data partitioning by date columns or other high cardinality keys, reducing the volume read in queries and optimizing cost and performance in Databricks.

Familiarity with clustering and Z-Ordering techniques in Databricks/Delta Lake for read optimization in high-volume tables.

Ability to identify and correct performance bottlenecks in SQL queries, such as avoiding full table scans, efficient use of joins, aggregations and window functions.

Ability to estimate and manage the impact of complex models on cloud processing costs, proposing solutions that balance performance and operational efficiency.

Requirements

Tools and Technologies

  • Mastery of DBT (dbt Core or dbt Cloud), including creation of models, tests, macros and documentation.
  • Experience with Databricks and Delta Lake, including features such as Time Travel, VACUUM and OPTIMIZE.
  • Advanced SQL skills for building complex queries and analytical modeling.
  • Knowledge of Python for automation, data transformation and development of scripts to support pipelines.
  • Experience with pipeline orchestration using Apache Airflow and Databricks.
  • Experience with code versioning using GitHub, including review flows and team collaboration.

Behavioral Skills

  • Critical thinking and autonomy to conduct complex projects with multiple stakeholders.
  • Ability to communicate clearly with non-technical business areas, translating analytical needs into data solutions.
  • Strong sense of ownership over the quality and reliability of the data delivered.

Differences

  • Experience with Unity Catalog or other data governance and cataloging tools.
  • Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know.
Role:
Analytics Engineer Sênior
Job Type:
Full Time

More jobs at Caju

Similar Data Engineer jobs at other companies