Exl logo

Exl

Agentic AI Data Engineer

Gurugram, Haryana, India

Job details
Agentic AI Data Engineer Agentic AI Data Engineer Role Overview Total Experience required : 5-10 Years We are seeking a highly skilled Agentic AI Data Engineer to design, build, and optimize intelligent, autonomous data systems that power next-generation AI applications. This role blends data engineering, machine learning infrastructure, and emerging agent-based AI frameworks to enable scalable, self-orchestrating pipelines and decision-making systems. You will work at the intersection of data platforms, large language models (LLMs), and cloud-native architectures—building systems that can reason, act, and adapt autonomously. Key Responsibilities - Design and implement agentic AI systems that autonomously orchestrate data workflows and decision pipelines - Build scalable data pipelines for structured and unstructured data (batch + real-time) - Develop and manage LLM-powered applications using retrieval-augmented generation (RAG), tool use, and multi-agent frameworks - Integrate AWS AI/ML services into production-grade architectures - Develop and optimize data lakes, warehouses, and lakehouse architectures - Build APIs and microservices to expose AI/ML capabilities - Ensure data quality, governance, and security across pipelines - Collaborate with data scientists, ML engineers, and product teams to deploy AI solutions - Implement monitoring, logging, and observability for AI agents and pipelines - Optimize cost and performance of cloud-based AI workloads Required Technical Skills Cloud & AWS Ecosystem - Strong experience with AWS services, including: - Amazon S3, Glue, Lambda, Step Functions - Amazon Redshift / Athena - Amazon SageMaker (training, deployment, pipelines) - Amazon Bedrock (foundation models, agents, knowledge bases) AI/ML & Agentic Systems - Experience with LLMs and generative AI systems - Hands-on with agent frameworks (e.g., multi-agent orchestration, tool calling, planning systems) - Familiarity with AgentCore / agent orchestration platforms - Understanding of RAG architectures, embeddings, and vector databases - Experience with model deployment, inference optimization, and prompt engineering Data Engineering - Strong proficiency in Python and SQL - Experience with ETL/ELT tools and frameworks - Distributed data processing (Spark, PySpark, or similar) - Streaming technologies (Kafka, Kinesis, or similar) - Data modeling and schema design Data & AI Infrastructure - Experience with vector databases (e.g., Pinecone, FAISS, OpenSearch) - Knowledge of data lakehouse architectures (Delta Lake, Iceberg, Hudi) - Containerization (Docker) and orchestration (Kubernetes) - CI/CD for ML and data pipelines Preferred Qualifications - Experience building autonomous AI agents for enterprise use cases - Knowledge of multi-agent collaboration systems and planning algorithms - Familiarity with LangChain, LlamaIndex, or similar frameworks - Experience with MLOps and LLMOps practices - Understanding of graph-based workflows and knowledge graphs - Exposure to real-time AI systems and event-driven architectures Soft Skills - Strong problem-solving and system design skills - Ability to work in fast-paced, evolving AI environments - Effective communication and cross-functional collaboration - Curiosity and adaptability to emerging AI technologies Education & Experience - Bachelor’s or Master’s degree in Computer Science, Engineering, or related field - 4+ years of experience in data engineering or ML engineering - Hands-on experience with production-grade AI/ML systems Nice-to-Have - Experience with reinforcement learning or planning systems - Background in distributed systems design - Contributions to open-source AI/data projects - Certifications in AWS (e.g., Solutions Architect, Machine Learning Specialty) What You’ll Build - Autonomous data pipelines that self-heal and optimize - AI agents capable of reasoning over enterprise data - Scalable LLM-powered applications integrated with business workflows - Intelligent systems that move beyond automation into decision-making Responsibilities Key Responsibilities - Design and implement agentic AI systems that autonomously orchestrate data workflows and decision pipelines - Build scalable data pipelines for structured and unstructured data (batch + real-time) - Develop and manage LLM-powered applications using retrieval-augmented generation (RAG), tool use, and multi-agent frameworks - Integrate AWS AI/ML services into production-grade architectures - Develop and optimize data lakes, warehouses, and lakehouse architectures - Build APIs and microservices to expose AI/ML capabilities - Ensure data quality, governance, and security across pipelines - Collaborate with data scientists, ML engineers, and product teams to deploy AI solutions - Implement monitoring, logging, and observability for AI agents and pipelines - Optimize cost and performance of cloud-based AI workloads Qualifications Required Technical Skills Cloud & AWS Ecosystem - Strong experience with AWS services, including: - Amazon S3, Glue, Lambda, Step Functions - Amazon Redshift / Athena - Amazon SageMaker (training, deployment, pipelines) - Amazon Bedrock (foundation models, agents, knowledge bases) AI/ML & Agentic Systems - Experience with LLMs and generative AI systems - Hands-on with agent frameworks (e.g., multi-agent orchestration, tool calling, planning systems) - Familiarity with AgentCore / agent orchestration platforms - Understanding of RAG architectures, embeddings, and vector databases - Experience with model deployment, inference optimization, and prompt engineering Data Engineering - Strong proficiency in Python and SQL - Experience with ETL/ELT tools and frameworks - Distributed data processing (Spark, PySpark, or similar) - Streaming technologies (Kafka, Kinesis, or similar) - Data modeling and schema design Data & AI Infrastructure - Experience with vector databases (e.g., Pinecone, FAISS, OpenSearch) - Knowledge of data lakehouse architectures (Delta Lake, Iceberg, Hudi) - Containerization (Docker) and orchestration (Kubernetes) - CI/CD for ML and data pipelines Preferred Qualifications - Experience building autonomous AI agents for enterprise use cases - Knowledge of multi-agent collaboration systems and planning algorithms - Familiarity with LangChain, LlamaIndex, or similar frameworks - Experience with MLOps and LLMOps practices - Understanding of graph-based workflows and knowledge graphs - Exposure to real-time AI systems and event-driven architectures