Job Overview
Job description
- Salary:
- CAD 75,000 per year
- Location:
- Canada
- Work arrangement:
- On-site
Role Summary
- Data Engineer (Mid-level)
- Location: Remote (Quebec/Canada)
- About Irth Solutions
Irth Solutions is a leading provider of cloud-based SaaS software for damage prevention, asset integrity, stakeholder engagement and land management, helping energy, utility, telecom, and infrastructure companies protect their critical network infrastructure. With nearly three decades of industry experience, Irth serves customers across North America and continues to expand its platform with new data-driven and AI-powered capabilities.
We are looking for a Data Engineer to design, build, and maintain data ingestion and processing pipelines in Databricks, transforming high-volume external data sources into clean, structured, reliable inputs for downstream analysis and intelligence. This role will primarily support our Stakeholder Engagement offering, working closely with our Data Scientist and application teams.
Responsibilities
- Data Pipeline Development (Primary Responsibility)
- Design, build, and maintain ingestion pipelines from high-volume external APIs, capable of running continuously and reliably at scale.
- Implement ingestion and transformation workflows using Databricks (Spark/PySpark, SQL, Delta Live Tables), applying medallion architecture patterns (Bronze → Silver → Gold) to move from raw ingested content to clean, structured, analysis-ready data.
- Build the infrastructure for deduplication and relevance filtering of incoming content, implementing filtering logic and quality criteria defined in collaboration with the Data Scientist.
- Implement schema evolution handling and data validation rules as data sources and formats change over time.
- Platform & Storage Implementation
- Configure and manage Delta Lake storage structures, tables, partitions, and optimization routines (OPTIMIZE, Z-ORDER, VACUUM).
- Design and evolve data schemas that balance query performance, cost, and maintainability as data volume grows.
- Maintain clear metadata and documentation of table structures to support easy consumption by the Data Science and application teams.
- Reliability, Monitoring & Operational Support
- Ensure pipeline reliability and observability: error handling, retries, monitoring, and alerting for a continuously running system.
- Adapt pipelines to evolving external API contracts, rate limits, authentication changes, and new data sources.
- Troubleshoot pipeline failures, perform recovery, and tune performance as needed.
- Orchestration & Automation
- Build, schedule, and monitor workflows using Databricks Workflows, Delta Live Tables, or similar orchestration tools.
- Contribute to CI/CD pipelines for code deployment, versioning, and environment management.
- Collaboration & Documentation
- Work closely with the Data Scientist to expose clean, well-structured data feeding LLM/NLP pipelines and downstream models.
- Participate in technical decisions around data architecture and propose structuring solutions as the team's needs evolve.
- Document pipelines, data dictionaries, job schedules, and transformation logic.
- Support the onboarding of new data sources and pipelines as the product expands to additional solution areas.
- Ingénieur(e) de données (Intermédiaire)
- Lieu: Télétravail (Québec/Canada)
À propos d' Irth Solutions
Irth Solutions est un fournisseur de premier plan de logiciels SaaS infonuagiques pour la prévention des dommages, l'intégrité des actifs, la mobilisation des parties prenantes et la gestion foncière, aidant les entreprises des secteurs de l'énergie, des services publics, des télécommunications et des infrastructures à protéger leurs réseaux d'infrastructure critiques. Avec près de trois décennies d'expérience dans l'industrie, Irth dessert des clients à travers l'Amérique du Nord et continue d'élargir sa plateforme avec de nouvelles capacités axées sur les données et l'intelligence artificielle.
À propos du poste
Nous sommes à la recherche d'un(e) ingénieur(e) de données pour concevoir, construire et maintenir des pipelines d'ingestion et de traitement de données dans Databricks, transformant des sources de données externes à haut volume en données propres, structurées et fiables pour l'analyse et l'intelligence en aval. Ce poste appuiera principalement notre offre de mobilisation des parties prenantes (Stakeholder Engagement), en collaboration étroite avec notre data scientist et les équipes de développement applicatif.
Responsabilités principales
- Développement de pipelines de données (responsabilité principale)
- Concevoir, construire et maintenir des pipelines d'ingestion à partir d'API externes à haut volume, capables de fonctionner de façon continue et fiable à grande échelle.
- Implement ingestion and transformation flows with Databricks (Spark/PySpark, SQL, Delta Live Tables), applying medallion architecture patterns (Bronze → Silver → Gold) to transform raw ingested content into clean, structured and analysis-ready data.
- Build the infrastructure for deduplication and relevance filtering of incoming content, implementing the filtering logic and quality criteria defined in collaboration with the data scientist.
- Implement management of evolving data validation schemas and rules as data sources and formats evolve.
- Platform and storage implementation
- Configure and manage Delta Lake storage structures, tables, partitioning and optimization routines (OPTIMIZE, Z-ORDER, VACUUM).
- Design and evolve data schemas that balance query performance, cost, and maintainability as data volume grows.
- Maintain clear documentation of metadata and table structures to facilitate their use by data science and application development teams.
- Reliability, monitoring and operational support
- Ensure the reliability and observability of pipelines: error management, recovery mechanisms (retries), monitoring and alerts for a continuously operating system.
- Adapt pipelines to changes in external API contracts, rate limits, authentication methods, and the addition of new data sources.
- Diagnose pipeline failures, carry out the necessary repairs and optimize performance if necessary.
- Orchestration and automation
- Build, schedule, and monitor workflows with Databricks Workflows, Delta Live Tables, or equivalent orchestration tools.
- Contribute to CI/CD pipelines for code deployment, versioning and environment management.
- Collaboration and documentation
- Collaborate closely with the data scientist to provide clean, well-structured data feeding into LLM/NLP pipelines and downstream models.
- Participate in technical decisions related to data architecture and propose structuring solutions as the team's needs evolve.
- Document pipelines, data dictionaries, execution schedules and transformation logic.
- Support the integration of new data sources and pipelines as the product expands into other solution areas.
Requirements
- Strong preference for candidates residing in Quebec, with mastery of French (spoken and written) being an important asset in addition to English.
- Strong preference for candidates residing in Quebec, with fluency in French (spoken and written) as a strong asset in addition to English
- 3 to 5 years of experience in data engineering, with solid experience building and operating production-grade data pipelines.
- Familiarity with data modeling, data quality, and schema evolution.
- Solid understanding of data pipeline reliability practices: monitoring, alerting, and handling failures gracefully in a continuously running system.
- Hands-on experience with Databricks (or an equivalent Spark-based environment): schema design, Delta Lake, performance tuning, and pipeline orchestration.
- Experience with at least one major cloud (Azure preferred; AWS/GCP also beneficial).
- Experience integrating with external APIs at scale: authentication, pagination, rate limiting, retries, error handling.
Strong proficiency in Python and SQL
Comfortable working with unstructured/semi-structured text data at scale.
Nice to have Qualifications
- LLM prompting experience and/or basic understanding of AI/NLP concepts
- Exposure to medallion architecture or lakehouse best practices.
- Experience with orchestration frameworks (ADF, Workflows, Airflow, DBX, etc.).
- Experience with CI/CD tools and version control (Git, GitHub Actions or equivalent).
- Basic understanding of security practices: RBAC, encryption, credential management.
- Databricks certification (Data Engineer Associate or equivalent).
AI Use in Hiring
- As part of our hiring process, this role may use artificial intelligence or automated tools to assist with reviewing and screening applications. These tools support, but do not replace, human judgment in making hiring decisions.
- 3-5 years of data engineering experience, with strong experience building and operating data pipelines in production.
- Knowledge of data modeling, data quality and schema evolution.
- Good understanding of data pipeline reliability practices: monitoring, alerting and fault management in a continuously operating system.
- Hands-on experience with Databricks (or equivalent Spark-based environment): schema design, Delta Lake, performance optimization and pipeline orchestration.
- Experience with at least one major cloud provider (Azure preferred; AWS/GCP also an asset).
- Experience in integrating large-scale external APIs: authentication, paging, rate limits, recovery mechanisms, error handling.
- Strong command of Python and SQL.
- Comfortable with processing large-scale unstructured or semi-structured text data.
Advantages
- Experience in prompt engineering with LLM and/or basic knowledge in AI/NLP.
- Knowledge of medallion architecture or lakehouse best practices.
- Experience with orchestration frameworks (ADF, Workflows, Airflow, DBX, etc.).
- Experience with CI/CD and version control tools (Git, GitHub Actions or equivalent).
- Basic knowledge of security practices: RBAC, encryption, identifier management.
- Databricks certification (Data Engineer Associate or equivalent).
Using AI in the recruitment process
As part of our recruitment process, this position may use artificial intelligence or automated tools to support the review and pre-screening of applications. These tools support human judgment in making hiring decisions, but do not replace it.
The salary range for this role is CAD $75,000
- This range reflects the base salary only and does not include any additional compensation components.
- Any offer of employment is dependent on several factors, including, but not limited to, the candidate’s experience, skills, qualifications, and location.
- The salary range for this position is between CAD $75,000
- This scale reflects base salary only and does not include any other compensation components. Any offer of employment depends on several factors, including, but not limited to, the applicant's experience, skills, qualifications and location.
Benefits
- Competitive Salary – A competitive compensation package based on experience and qualifications.
- Medical, Dental, and Vision Insurance – Comprehensive insurance coverage to support you and your family.
- 401(k) Plan with Company Match.
- Generous Paid Time Off (PTO) – Time off to support work-life balance and personal needs.
- Company-Paid Holidays – Paid holidays throughout the year.
- Flexible Work Options – Work-from-home opportunities are available, depending on role and business needs.
- On-Call Compensation – Additional pay for eligible on-call shifts.
- Competitive Salary — Competitive compensation based on experience and qualifications.
- Medical, Dental and Vision Insurance — Comprehensive insurance coverage for you and your family.
- Retirement plan with employer contribution.
- Generous Paid Time Off — Time off to promote work-life balance.
- Company Paid Holidays — Paid time off throughout the year.
- Flexible working arrangements — Possibility of teleworking depending on the position and the needs of the company.
- Compensation for availability — Additional remuneration for eligible on-call shifts.
- Role:
- Data Engineer
- Job Type:
- Full Time