Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Foundations of Multimodal LLMs on Vertex AI
- An overview of multimodal capabilities available within Vertex AI.
- Exploration of Gemini models and the specific modalities they support.
- Key enterprise and research use cases for multimodal AI.
Preparing the Development Environment
- Configuring Vertex AI specifically for multimodal workflow execution.
- Techniques for managing datasets that span different modalities.
- Practical Lab: Setting up the environment and preparing multimodal datasets.
Long Context Windows and Advanced Reasoning
- Concepts and mechanics behind long-context workflows.
- Applications in strategic planning and complex decision-making processes.
- Practical Lab: Implementing and testing long-context analysis techniques.
Designing Cross-Modal Workflows
- Strategies for combining text, audio, and image analysis in a single flow.
- Methods for chaining multimodal processing steps within pipelines.
- Practical Lab: Designing and building a functional multimodal pipeline.
Optimizing Gemini API Parameters
- Techniques for configuring multimodal inputs and outputs effectively.
- Best practices for optimizing inference speed and overall efficiency.
- Practical Lab: Tuning Gemini API parameters for specific performance goals.
Advanced Applications and System Integrations
- Developing interactive multimodal agents and intelligent assistants.
- Methods for integrating external APIs and third-party tools.
- Practical Lab: Building a complete multimodal application from scratch.
Evaluation Strategies and Iterative Improvement
- Approaches for testing multimodal model performance.
- Key metrics for assessing accuracy, alignment, and data drift.
- Practical Lab: Evaluating and refining multimodal workflow outputs.
Course Summary and Recommended Next Steps
Requirements
- Proficiency in Python programming.
- Hands-on experience with developing machine learning models.
- Working familiarity with multimodal data types, including text, audio, and images.
Intended Audience
- AI Researchers.
- Advanced Developers.
- Machine Learning Scientists.
14 Hours