Complir is building an AI platform for product compliance, turning global regulatory text into structured, queryable data. As Data Engineer, you’ll own the crawling, parsing, and change-detection pipelines that keep the platform’s regulatory data current and provably correct. You’ll work on-site in Copenhagen alongside founders, AI engineers, and compliance experts.
Responsibilities
Build and operate crawling and monitoring across hundreds of regulatory sources.
Parse raw legal text (HTML, PDFs, scans) into a normalized requirement schema using LLM-assisted extraction.
Build change detection that catches new acts, amendments, and repeals the day they land and triggers downstream updates.
Own data quality end to end: provenance, versioning, and coverage metrics.
Shape the architecture, tooling, and data culture as the team grows.
Requirements
Several years building data pipelines, crawling infrastructure, or large-scale extraction systems.
Experience with or strong interest in LLM-assisted extraction: structured outputs, evals, and failure modes of models reading dense text.
Treat data correctness as sacred — provenance, versioning, and traceability.
Comfortable with messy sources: broken HTML, scanned PDFs, and multilingual legal prose.
Product-minded and able to thrive in a fast-moving, small-team environment.
Complir is an AI-powered software company that provides infrastructure for product compliance, specifically for global retailers. The platform automates complex workflows such as mapping product data to regulatory requirements, generating documentation, and tracking risk across various markets.