AI Instructor Live Labs Included

OpenAI: Multimodal Applications with GPT-4o

Build multimodal AI applications with GPT-4o vision, Whisper audio transcription, TTS speech synthesis, and gpt-image-1 image generation.

Intermediate
4h 40m
10 Lessons
OPENAI-203
OpenAI Multimodal Developer Badge

View badge details

About This Course

Build real-world applications that combine vision, audio, and image generation using the OpenAI API. Learn to analyze images and documents with GPT-4o, transcribe and translate audio with Whisper, generate expressive speech with TTS, create and edit images with gpt-image-1, and orchestrate all modalities in a multimodal meeting assistant capstone project.

Course Curriculum

10 Lessons
01
AI Lesson
AI Lesson

Vision: Images & Document Understanding

20m

Learn how GPT-4o processes images using multi-modal inputs. Covers URL-based and base64 image encoding, detail levels (low/high) and their token impact, and structured output extraction from visual content like invoices and documents.

02
Lab Exercise
Lab Exercise

Vision Applications

30m 4 Exercises

Practice using GPT-4o vision API by analyzing images via URL and base64 encoding, comparing detail levels, and extracting structured data from invoice images using Pydantic models.

03
AI Lesson
AI Lesson

Audio Input & Speech Recognition with Whisper

20m

Learn to transcribe and translate audio using OpenAI Whisper. Covers basic transcription, word-level timestamps, translation to English, and building voice command routing pipelines by combining Whisper with GPT classification.

04
Lab Exercise
Lab Exercise

Speech Recognition Pipeline

30m 4 Exercises

Practice building audio transcription workflows with Whisper: basic transcription, word-level timestamps, multilingual translation, and a voice command routing system powered by Whisper + GPT classification.

05
AI Lesson
AI Lesson

Audio Output: Text-to-Speech

20m

Learn to generate high-quality speech with OpenAI TTS models. Covers standard synthesis with tts-1, all six built-in voices, expressive speech with gpt-4o-mini-tts style instructions, and low-latency audio streaming.

06
Lab Exercise
Lab Exercise

Voice Output Applications

30m 4 Exercises

Practice generating speech with OpenAI TTS: synthesize audio with tts-1, compare all six voices, create expressive speech with style instructions using gpt-4o-mini-tts, and implement low-latency audio streaming.

07
AI Lesson
AI Lesson

Image Generation with gpt-image-1

20m

Learn to generate and edit images using OpenAI's gpt-image-1 model. Covers text-to-image generation, quality and size settings, inpainting with masks, and batch generation of multiple variations in a single API call.

08
Lab Exercise
Lab Exercise

Image Generation Applications

30m 4 Exercises

Practice generating and editing images with gpt-image-1: create images from text prompts, compare quality settings, perform inpainting edits with masks, and generate multiple image variations in a single API call.

09
AI Lesson
AI Lesson

Capstone Briefing: Multimodal Meeting Assistant

20m

Architecture overview of the capstone project: a meeting assistant that transcribes audio with Whisper, analyzes slides with GPT-4o vision, generates structured reports with GPT-4.1, and narrates summaries with TTS — demonstrating orchestration of all multimodal capabilities.

10
Lab Exercise
Lab Exercise

Capstone Project: Multimodal Meeting Assistant

1h 0m 5 Exercises

Build a complete multimodal meeting assistant that transcribes audio with Whisper, analyzes presentation slides with GPT-4o, generates a structured JSON report with GPT-4.1, and narrates the executive summary with expressive TTS.

This course includes:

  • 24/7 AI Instructor Support
  • Live Lab Environments
  • 5 Hands-on Lessons
  • Completion Badge
OpenAI Multimodal Developer Badge

Earn Your Badge

Complete all lessons to unlock the OpenAI Multimodal Developer achievement badge.

Skill Level Intermediate
Total Duration 4h 40m
OpenAI Multimodal Developer Badge
Achievement Badge

OpenAI Multimodal Developer

Awarded for completing Multimodal Applications with GPT-4o. Demonstrates ability to analyze images with GPT-4o vision, transcribe and translate audio with Whisper, generate expressive speech with TTS, create and edit images with gpt-image-1, and orchestrate multimodal pipelines.

Course OpenAI: Multimodal Applications with GPT-4o
Criteria Complete all lessons and exercises in OPENAI-203: Multimodal Applications with GPT-4o
Valid For 730 days

Skills You'll Earn

GPT-4o Vision Whisper Text-to-Speech Image Generation Multimodal AI Pipeline Orchestration

Complete all lessons in this course to earn this badge