Jobiglo

No results.

This job is no longer available

This job expired on 11/08/2026. It no longer accepts applications.

Data Engineer – AI/LLM Delivery Unit

Innodata Inc.

Mid 🇬🇧 English
Python Scala SQL Airflow Dagster Snowflake BigQuery Redshift Spark Kafka Kinesis Pinecone FAISS Weaviate AWS Azure GCP Docker CI/CD Terraform

Job description

About the role

We are seeking a Data Engineer to join our AI/LLM Delivery Unit, responsible for building scalable data pipelines and infrastructure that power AI and machine learning solutions. The role will enable LLM‑based applications, data workflows, and AI model lifecycle management across structured and unstructured datasets.

Key responsibilities

  • Design, build and maintain scalable ETL/ELT pipelines for both structured and unstructured data, ensuring reliable ingestion, transformation and delivery.
  • Support AI/ML and LLM workflows, including training, fine‑tuning and evaluation datasets, and develop pipelines for text corpora, embeddings, vector stores and Retrieval‑Augmented Generation systems.
  • Automate data extraction, transformation and validation, implementing batch and real‑time processing solutions to improve operational efficiency.
  • Implement data quality, validation, monitoring and governance processes, maintaining data integrity, lineage and compliance with security standards.
  • Collaborate closely with Data Scientists, ML Engineers and delivery teams to translate business and AI requirements into scalable data architectures.

Required profile

  • Bachelor’s degree in Computer Science, Data Engineering, Information Systems or related field (advanced degree a plus).
  • 3–7+ years of experience in data engineering or related roles, preferably supporting AI/ML or analytics platforms.
  • Demonstrated experience with AI/LLM‑related data pipelines is a strong advantage.

Required skills

  • Programming: Python and/or Scala.
  • SQL and database design.
  • ETL orchestration tools such as Airflow or Dagster.
  • Data warehouses: Snowflake, BigQuery, Redshift.
  • Distributed processing: Apache Spark.
  • Streaming: Kafka, Kinesis (optional).
  • Unstructured data pipelines, NLP datasets, embeddings and vector databases (Pinecone, FAISS, Weaviate).
  • Cloud platforms: AWS, Azure, GCP.
  • Containerisation and DevOps: Docker, CI/CD pipelines, Infrastructure‑as‑Code (Terraform a plus).

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Innodata Inc..
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Why are you reporting this job?

Thank you for your report. We will review this job.

A question about this job?

Ask it here: you will get the full job summary by e-mail, right away.

💬 Chat with us on Telegram

Published 3 months ago

10 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

Innodata Inc.