Pinnipedia, a Berlin startup, is looking for a Senior AI/Knowledge Graph Engineer to build a cloud platform that automates the creation of audit-ready IT-security concepts. The role involves owning the end-to-end pipeline that turns unstructured documents into a validated, queryable knowledge graph using LLM extraction agents, graph databases, and evaluation frameworks.
Responsibilities
Build LLM extraction pipelines for document chunking, property/relationship extraction, cross-chunk reconciliation, and gap detection.
Design knowledge graph schema as typed Pydantic models, including Cypher access patterns and indexing strategy, and manage schema evolution.
Develop deterministic rule engines for cases where code outperforms LLM judgment, maintaining clear contracts between deterministic and probabilistic components.
Ensure data validation and quality through schema enforcement, required-property contracts, audit trails, and eval harnesses (including LLM-as-judge).
Handle live data operations: backfills, coordinated migrations across relational and graph stores, observability on extraction throughput/quality, and incident response.
Requirements
5+ years shipping data/AI systems to production with real customers, including on-call experience for live pipelines.
Strong Python (typed, modern) and SQL, with comfort using PostgreSQL under load.
Hotjar, now part of Contentsquare, provides an experience intelligence platform offering product experience insights. Its core tools include heatmaps, session recordings, surveys, and feedback to help teams understand user behavior and improve website or app experiences.
Production experience with at least one graph database (Neo4j preferred; Neptune, ArangoDB, TigerGraph acceptable) including schema design and query tuning.
Production LLM pipeline experience: structured output, agent orchestration, prompt/version management, and evaluation frameworks (PydanticAI, LangChain, DSPy, or Instructor).
Durable workflow orchestration in production (DBOS, Temporal, Airflow, Prefect, Dagster).
Test-first discipline with integration tests against real datastores (Testcontainers or equivalent).
Fluent English skills.
Nice to Have
Experience with regulated, compliance-driven, or standards-heavy extraction domains (legal, medical, financial, security/audit).
Experience designing deterministic evaluators alongside LLM components and knowing when to use which.
Contributions to data contracts, schema governance, or ontology work.