Job Overview
Job description
- Location:
- Colombia
- Work arrangement:
- On-site
Role Summary
Robots & Pencils is an applied AI engineering firm building the next frontier of business architecture. We design and ship AI co-workers that integrate into enterprise operations and deliver measurable results for our clients. We're all in on AWS, combining deep UX capability with senior engineering talent to get AI into production fast and keep it there.
We’ve earned the trust of leaders across Consumer Products and Retail, Education, Energy, Financial Services, Healthcare, and Manufacturing and more, and earned a reputation as the nimble alternative to traditional global systems integrators. Founded in 2009, with delivery centers in Canada, the United States, Eastern Europe, and Latin America, we are smaller, faster, and more senior by design. Our teams average 15+ years of experience. We move fast, sweat the details, and build things that actually ship.
We’re looking for a Senior Testing Engineer to lead quality across complex software features, AI/ML systems, and integrated platforms. This role is ideal for an experienced engineer who is knowledgeable across both automation and AI testing, enjoys hands-on work, and is growing into owning testing strategy and architecture.
In this role, you will work as part of a cross-functional team, designing automation frameworks, validating ML models and AI workflows, and driving quality throughout the delivery process. You’ll collaborate closely with developers, ML engineers, and product teams to deliver high-quality features and develop the technical instincts that come from shipping things that matter.
Why This Role Matters
At Robots & Pencils, we design AI systems for a human world. Our name says it all. Robots and pencils means engineering paired with creativity, because every agent we ship has to work for real people in real workflows. That balance is baked into how we operate.
Every role here contributes directly to that mission. Here, you shape how AI systems integrate into enterprise operations, how teams move at real velocity, and how products create measurable impact for clients and the people they serve. We ship production-ready AI in 30 to 45 days. That pace demands people who take ownership, lead with craft, and care deeply about what they put their name on.
Responsibilities
- Craft & Delivery
- Design and develop scalable automation frameworks for UI, API, and integration testing (e.g., Cypress, Playwright, Selenium, PyTest, TestNG, Cucumber)
- Integrate automated tests into CI/CD pipelines and drive CI/CD quality gates (e.g., GitHub Actions, GitLab CI, Jenkins)
- Design AI-specific test strategies and validate ML models across key metrics including accuracy, precision, recall, bias, and drift
- Test data pipelines, feature engineering processes, and validate LLM responses for hallucination risk and output consistency
- Perform test planning, review automation code, and ensure best practices are applied across the test codebase
- Monitor model quality in production environments and support release validation and quality assurance processes
- Bring an AI-forward mindset to your daily work, using tools like Claude, Cursor, and other modern AI assistants to ship higher-quality work at pace
- Collaboration & Communication
- Collaborate closely with developers, ML engineers, DevOps, and product teams across the full SDLC
- Communicate quality status, testing findings, and risks clearly to stakeholders
- Participate actively in sprint ceremonies, design reviews, and release planning
- Leadership & Influence
- Lead quality efforts end-to-end with growing ownership of testing strategy and automation architecture
- Contribute to QA standards and best practices, and identify opportunities to improve coverage, reliability, and automation
- Begin mentoring junior engineers, sharing knowledge and supporting their growth
Requirements
- 3–4+ years of experience in QA, data QA, or AI testing, with solid knowledge across both automation and AI testing
- Strong programming skills in Python and at least one other language (e.g., Java, JavaScript)
- Hands-on experience designing and maintaining automation frameworks (e.g., Selenium, Cypress, Playwright, PyTest, TestNG, Cucumber)
- Strong understanding of the ML lifecycle and experience with model evaluation metrics (accuracy, precision, recall, bias, drift)
- Knowledge of AI testing methodologies including LLM validation, hallucination testing, and data pipeline validation
- Experience with CI/CD pipelines, API automation, and version control (e.g., GitHub Actions, Postman, Git)
- Familiarity with MLOps pipelines and model monitoring in production environments
- Strong understanding of Agile methodologies and experience with test management tools (e.g., Jira, TestRail)
- Demonstrable usage of AI-forward tools such as Claude and Cursor
- Experience with mobile automation, performance testing, LLM testing frameworks, prompt engineering, or cloud ML platforms is a plus (e.g., Appium, JMeter, AWS SageMaker, Azure ML)
- Helpful Extras and Unique Skills
You’ll Do Well Here if You Are
- A doer. You see something broken and fix it. You'd rather move on clarity than wait for certainty.
- A fast learner who knows you don't know everything. The AI landscape changes weekly. You're senior enough to know better and curious enough to keep learning anyway.
- Direct in a way that makes the work better. You give honest feedback. You'd rather have the hard conversation than blow smoke.
- Obsessed with craft. You know genius is in the details. You ship exceptional, not perfect, and you don't put your name on work you wouldn't stand behind.
- Built for ownership. You honor commitments, admit mistakes fast, and back your teammates when a decision costs something. No handoffs, no finger-pointing.
- All in. You treat clients' businesses like your own. You take the work seriously without taking yourself seriously.
- Resourceful when the budget, timeline, or team is tight. Constraints don't slow you down. They sharpen you.
- Glad to be in the room with people who care as much as you do. Our teams average fifteen-plus years of experience. We hire people who push each other to do better work.
About the Company
Why Join Robots & Pencils?
At Robots & Pencils, we build intelligent systems for real-world environments, blending creativity, engineering excellence, and applied AI to help organizations work smarter. Our teams partner directly with leading North American organizations to design and deliver production-ready digital and AI solutions that create measurable impact.
As an AI Engineer, you will build production AI systems alongside experienced architects and senior engineers, strengthening your foundation in machine learning, scalable system design, and modern cloud-native delivery. You’ll collaborate closely with US-based clients, gaining hands-on exposure to complex problem-solving while contributing to meaningful, real-world outcomes.
By joining Robots & Pencils, you also benefit from our Advanced AWS Partnership and participation in the highly exclusive AWS Patterns Partnership, a distinction held by only 11 companies worldwide out of more than 190,000 AWS partners. This provides access to specialized training, early access to emerging cloud capabilities, and advanced engineering resources that support accelerated professional growth.
This is a role for builders who want to sharpen their craft, grow quickly, and deliver impactful AI systems within an elite global engineering ecosystem.
- Role:
- Senior Testing Engineer
- Job Type:
- Full Time
Company profile

Robots & Pencils
robotsandpencils.comRobots & Pencils is an applied AI engineering firm that builds generative and agentic AI systems on Amazon Web Services. It is an AWS Advanced Tier Services Partner and one of 11 inaugural AWS Pattern Partners. The company builds agentic systems on Amazon Bedrock and Amazon Bedrock AgentCore, and works with North American organizations to design and deliver production digital and AI solutions, taking projects from pilot to live use.