The tutorial (all code, which is refreshing).
1️⃣ RAG
2️⃣ Evaluation and metrics
3️⃣ Reranking
4️⃣ Finetuning embeddings
5️⃣ Finetuning LLMs
The image shows a deployment on Google Vertex (Apache Airflow) pipelines.

ZenML is an open-source framework designed to simplify the creation of Machine Learning (ML) pipelines. It helps engineers and data scientists manage the lifecycle of ML models, from experimentation to deployment, in a structured and reproducible way.
Key Features of ZenML
- Pipeline-Oriented Design: ZenML encourages you to break your ML workflows into modular, reusable, and composable steps, such as data preprocessing, training, evaluation, and deployment.
- Flexibility with Integrations:
- ZenML integrates with various tools and frameworks, such as TensorFlow, PyTorch, Scikit-learn, and others.
- It supports deployment solutions like Kubernetes, AWS SageMaker, and others.
- Integration with experiment tracking tools (e.g., MLflow, Weights & Biases).
- Version Control for ML Pipelines:
- Tracks artifacts and metadata throughout the ML workflow.
- Ensures reproducibility by versioning pipelines and their dependencies.
- MLOps-Ready:
- ZenML bridges the gap between data science and production by enabling seamless collaboration between teams.
- It supports continuous integration and deployment (CI/CD) practices for ML workflows.
- Extensible and User-Friendly:
- Developers can extend ZenML with custom steps, pipelines, and integrations.
- It provides a clean and intuitive API for defining ML pipelines.
Use Cases of ZenML
- Building end-to-end ML workflows.
- Scaling ML pipelines with cloud infrastructure.
- Monitoring and managing deployed models.
- Creating reproducible experiments for ML research.
ZenML provides a robust framework for teams seeking to adopt MLOps practices, focusing on simplicity, scalability, and adaptability to various tech stacks.
What is RAG?
RAG is an approach that combines:
- Retrieval: Fetching relevant documents or data from a knowledge base (e.g., databases, document stores, or vector embeddings).
- Augmentation: Feeding the retrieved data as context to a language model (like GPT) for generating more informed and accurate responses.
RAG is commonly used in tasks like:
- Question answering systems.
- Document summarization.
- Conversational AI with domain-specific knowledge.
How ZenML Helps in RAG Pipelines
ZenML facilitates building and managing the complex pipeline components involved in a RAG system. Here’s how it can contribute:
1. Pipeline Management
ZenML helps structure RAG systems into modular, reusable steps:
- Data Ingestion and Preprocessing: Fetching and preparing data to create or update the knowledge base.
- Embedding Generation: Generating vector embeddings using models like Sentence Transformers or OpenAI embeddings.
- Indexing: Creating a searchable index (e.g., using tools like FAISS, Pinecone, or Weaviate).
- Retrieval Step: Setting up retrieval logic to fetch relevant documents based on a query.
- Augmentation and Generation: Integrating with language models (e.g., GPT) to use the retrieved documents for response generation.
- Evaluation and Feedback: Adding a feedback loop to evaluate the system’s performance and retrain if necessary.
2. Integrations for RAG Components
ZenML can integrate with the tools commonly used in RAG:
- Vector Stores: FAISS, Pinecone, Weaviate.
- Embedding Models: Hugging Face, OpenAI, or custom models.
- Language Models: GPT-based models or open-source alternatives like LLaMA or Bloom.
- Experiment Tracking: MLflow or Weights & Biases to monitor pipeline performance.
- Deployment Tools: Kubernetes or cloud services like AWS, GCP, and Azure.
3. Reproducibility
ZenML ensures that the pipeline components (e.g., embedding generation, indexing) are version-controlled, making it easy to:
- Reproduce the RAG system’s results.
- Debug or refine individual components of the pipeline.
4. Scalability and Automation
ZenML supports scaling RAG pipelines by:
- Running pipelines on distributed infrastructure.
- Automating updates to the knowledge base or retraining models when new data is added.
5. MLOps for RAG
ZenML provides out-of-the-box MLOps capabilities to:
- Monitor the retrieval and generation performance.
- Automate retraining or updating embeddings as the knowledge base evolves.
- Manage the lifecycle of deployed RAG systems, ensuring they remain up-to-date.
Example: RAG Workflow with ZenML
- Step 1: Fetch data from a document repository or API.
- Step 2: Process and clean the data.
- Step 3: Generate embeddings for the documents and store them in a vector database (e.g., FAISS or Pinecone).
- Step 4: Set up a retrieval mechanism to fetch top-k relevant documents for a given query.
- Step 5: Use the retrieved documents as input to a language model for generating responses.
- Step 6: Evaluate the system’s responses and retrain as necessary.
ZenML provides the tools to orchestrate, track, and manage this entire workflow seamlessly, ensuring a robust and scalable RAG system.