Senior AI Engineer, Data Quality & Pipeline Automation
- Pay
- Pay not listed
- Where
- Remote
- Type
- Contract
- Eligibility
- Country eligibility not stated. Check with the provider before applying.
- Listing
- Last verified Oct 6, 2026
Apply opens the provider’s application page. You apply there yourself: we never submit applications, and opening the page doesn’t mean you’ve applied. Nobody can promise you’ll be hired.
Some Apply links are referral links. If you apply through them, the provider may pay us a commission. Referral disclosure
Overview
Build and deploy an intelligent, autonomous operations agent that supervises data-pipeline infrastructure around the clock: predicting delivery delays before they happen, checking data quality against adaptive statistical baselines, and drafting code-level fixes for routine failures.
You own the agent end to end, from design and evaluation to production reliability.
Responsibilities
- Predictive health and SLA forecasting: models that compare in-flight pipeline progress, execution velocity and resource use with history to forecast SLA breaches 1–2 hours ahead and trigger re-queueing or escalation.
- Bottleneck profiling and self-healing: let the agent inspect execution DAGs and query profiles, isolate root causes, and draft fixes on isolated branches for syntax errors, upstream schema changes and column renames, each verified by dry-run assertions and submitted with test results for human approval.
- Statistical and semantic data-quality checks: rolling time-series baselines that handle seasonality, weekday swings and business close cycles, plus profiling of categorical entropy, column distributions and foreign-key orphan rates.
- Lineage reconciliation and quarantine: trace data from raw ingestion to reporting views, run cross-tier checksums and consistency checks before reporting cycles, and quarantine partitions with severe violations.
- Conversational copilot and incident management: root-cause incident briefs, and a natural-language interface for checking pipeline health, re-running partitions and analysing data distributions.
Requirements
- Strong production-grade Python for automation, agent logic, testing and data processing.
- Advanced SQL for profiling, validating and troubleshooting data in warehouses such as Snowflake, BigQuery or Databricks.
- Hands-on experience building and operating pipelines with Airflow (or Dagster/Prefect) and dbt, and understanding how failures start and spread.
- Data-quality and observability tools such as Great Expectations, Soda or Monte Carlo, with a solid grasp of anomaly detection, time-series baselines and forecasting.
- LLM and agent engineering: tool calling, multi-step agents and frameworks such as the Claude Agent SDK or LangGraph, with a focus on safety, evaluation and reliability.
Engagement and pay
- Pay: not stated. The listing says compensation is at market rate and asks you to give a specific hourly rate expectation.
- 40 hours per week, with at least 6 hours per day in IST.
- Independent contractor, approximately 2 months, starting November 1, 2026.
- Process: an AI interview (about 25 minutes), a practical code and AI evaluation exercise (about 30 minutes) focused on reviewing AI-generated code rather than algorithm puzzles, and a hiring-manager interview (about 20 minutes).
Eligibility
- Requires at least 6 hours per day in IST working hours.
- Country or residence limits: not stated in the listing.