AI Data Engineer

Responsibilities

AI Data Architecture & Platform Support

  • Architect and optimize high-throughput data pipelines and unstructured data processing frameworks to directly power the enterprise AI Foundation Platform.
  • Co-own the vector infrastructure alongside the AI Engineer Specialist, ensuring optimized embedding storage, indexing strategies, and fast vector search retrieval.
  • Build automated ETL/ELT pipelines capable of transforming raw enterprise data into AI-ready formats (chunking, metadata tagging, and cleaning).

Data Governance, Metadata & Cataloging

  • Implement metadata pipelines for comprehensive data catalog and data dictionary services, including schema capture, lineage tracking, and audit logging.
  • Enforce data governance, access controls, and privacy compliance across all data stores used for training, fine-tuning, and RAG retrieval.
  • Design and maintain operational data lineage to ensure transparency and auditability of data flowing into production AI models.

Advanced AI Component Integration

  • Develop and tune robust data ingestion pipelines for OCR, NLP processing, recommendation systems, and multi-modal AI inputs.
  • Optimize data tiering and caching strategies to reduce latency and infrastructure costs for live enterprise AI workloads.
  • Build secure, scalable data connectors linking legacy enterprise databases with modern LLM frameworks and autonomous agents.

Collaboration & Data Excellence

  • Partner directly with the AI Engineer Specialist to ensure seamless data delivery for real-world, high-concurrency production deployments.
  • Bridge traditional data engineering practices with modern MLOps, setting up automated data validation, quality checking, and drift detection at the data layer.
  • Provide data-level expertise during architecture reviews, ensuring scalable data design patterns across the entire AI project lifecycle.

Key Impacts

  • Ensure high-quality, AI-ready data ingestion at scale, dramatically reducing the time-to-market for enterprise AI assets.
  • Establish a secure and compliant data foundation that guarantees enterprise data privacy during AI model interactions.
  • Eliminate data bottlenecks in real-world production environments, ensuring sub-second latency for enterprise RAG and search systems.
  • Transition traditional data warehouses and lakes into future-ready, graph- and vector-enabled AI data infrastructure.

Qualification

Education & Experience

  • Master’s or Bachelor’s degree in Computer Engineering, Computer Science, Data Engineering, or a related technical field.
  • 5+ years of experience in heavy-duty Data Engineering, Big Data, or Distributed Systems, with strong production experience in an AI/ML context.
  • Proven track record of building large-scale data infrastructure supporting live, mission-critical applications.

Technical & AI Data Framework Expertise

  • Deep experience in Python and SQL, alongside modern big data tools (e.g., Spark, Kafka, or cloud equivalents).
  • Hands-on expertise in Vector Databases and unstructured data processing.
  • Strong understanding of data chunking strategies, text extraction (OCR/NLP pipelines), and metadata orchestration.
  • Familiarity with cloud data platforms (AWS/Azure/GCP) including managed data lakes, data warehouses, and cloud-native security/encryption.
  • Real-world experience implementing data catalogs, schema evolution, and automated data quality monitoring.

Leadership & Soft Skills

  • Excellent technical communication skills to align data pipeline designs with AI model engineering requirements.
  • Strong analytical and problem-solving mindset, comfortable handling noisy, unstructured enterprise data under tight deadlines.
  • Good verbal and written English communication skills for vendor engagement and cross-functional team collaboration.

We use cookies to enhance your site experience. by continuing to browse, you agree to our use of cookies & Cookies Policy  Click Settings to manage.

Privacy Preferences

You can set up your cookies preference by clicking the available sliders to ‘On’ or ‘Off’, except only Necessary cookies. Then, click “Save Preference”.

ยอมรับทั้งหมด
Manage Consent Preferences
  • Necessary cookies
    Always Active

    These cookies are necessary for the website to function and cannot be disabled. We use necessary cookies to enable core functionality such as security and network management, and, to allow you to browse the website normally. Without this cookies, the website would not be able to work properly.
    Cookies Details

  • Analytics cookies

    These cookies are used to collect information about how visitors use our website. We use the information to measure and improve the performance of our website. If you disable these cookies, we will not be able to use the information for improving our website.

  • Advertising Cookies

    These cookies are used to collect information about your activities (ex. sites or contents you’ve visited) to analyze and display content or advertisement that are relevant to your interests. If you disable these cookies, you will still see generic advertisement on your browser (not targeted content or advertisement).

  • Functional cookies

    These cookies enable the website to remember the information that you’ve pre-filled on the website (i.e. join our team or leave your contact for your interest in our product). The intention is to allow you to use the website more convenient. If you disable these cookies, then some or all of these services may not function properly.

Save