AI Instructor Live Labs Included

Structured Output + JSON with Claude

Produce schema-valid, downstream-ready structured output from Claude. Use tool calling as a JSON forcer, pair it with Pydantic and JSON Schema, build extraction pipelines for emails, PDFs, and HTML, stream structured outputs, and recover gracefully when validation fails.

Intermediate
10h 45m
10 Lessons
CLD-AI-108

About This Course

Deep dive on producing schema-valid, downstream-ready structured output from Claude. Cover tool_use as JSON forcer (from CLD-AI-102 L7), Pydantic + JSON Schema patterns, extraction pipelines for real-world documents (emails, PDFs, HTML), streaming structured outputs, and error recovery patterns when validation fails.

Course Curriculum

10 Lessons
01
AI Lesson
AI Lesson

Structured output overview: three paths

1h 0m

By the end of this lesson you will know the three ways to force Claude into structured JSON output and when each fits: prompt-and-parse (fragile, ~85%, useful only for prototypes), XML tags + parse (99%, good for shallow shapes, taught in CLD-AI-102), and tool_use with a JSON Schema (100% schema-valid, the production default). Anchors the rest of CLD-AI-108 on why tool_use is the shape you'll reach for by default. Sets up L2 hands-on where you extract emails via tool_use.

02
Lab Exercise
Lab Exercise

Email extraction pipeline - Lab Exercises

1h 10m 1 Exercises

By the end of this hands-on lab you will have built an email extraction pipeline that pulls sender, intent, urgency, deadline, and action_items out of five real-shape emails using Claude's tool_use forced-schema pattern with a JSON Schema — every response validates against the schema before your downstream code sees it. Uses the PythonAI container + Claude API proxy. Introduces the tool_use-as-extractor pattern you'll reuse in every remaining CLD-AI-108 lab.

03
AI Lesson
AI Lesson

Complex nested schemas + JSON Schema features

1h 0m

By the end of this lesson you will know how to shape production JSON Schemas Claude fills reliably — deeply nested objects, typed arrays of objects, discriminated unions via oneOf, regex-constrained string patterns, min/max bounds on numbers, and additionalProperties: false to prevent hallucinated fields — plus which schema shapes cause Claude to slip and how to rewrite them for higher reliability. Sets up L4 hands-on where you extract a hierarchical invoice into a nested schema.

04
Lab Exercise
Lab Exercise

Nested invoice extraction - Lab Exercises

1h 5m 1 Exercises

By the end of this hands-on lab you will have extracted a hierarchical Orion Analytics invoice into a nested JSON record: a customer object (name, tier, address) plus a line_items array with per-item type discrimination (product / service / discount) and computed total — all through Claude tool_use with a JSON Schema that uses oneOf discriminators + minItems bounds. Uses the PythonAI container + Claude API proxy. Proves complex nested schemas are still reliable when shaped correctly.

05
AI Lesson
AI Lesson

Extraction pipelines at scale

1h 0m

By the end of this lesson you will know how to scale extraction from one-shot to production volume — ThreadPoolExecutor concurrency to parallelize independent extractions, exponential-backoff retry on 429s and 529s, schema validation as the pass/fail gate, and writing extracted records to a warehouse or event stream. Anchors the mechanics you'll use in L6 hands-on to batch-extract 10 tickets in parallel with retry, and reappears in CLD-AI-111 for cost-optimization.

06
Lab Exercise
Lab Exercise

Batch extractor with concurrency + retry - Lab Exercises

1h 10m 1 Exercises

By the end of this hands-on lab you will have built a batch extractor that processes 10 Orion support tickets concurrently through ThreadPoolExecutor, retries transient failures (429 rate limit, 529 overload) with exponential-backoff-plus-jitter, and validates each returned record against the tool_use schema before it reaches downstream code. Uses the PythonAI container + Claude API proxy — measures throughput vs serial baseline and confirms retries fire on injected 429s.

07
AI Lesson
AI Lesson

Streaming structured output

1h 0m

By the end of this lesson you will know how to stream Claude tool_use output progressively — how content_block_delta events for a tool_use block carry partial JSON that you can parse field-by-field as it's generated, when streaming structured output improves perceived UI latency (rendering fields into a form as they arrive), and the parsing challenges (partial JSON is invalid JSON, so you use a streaming JSON parser). Sets up L8 hands-on where you stream tool_use into a live UI.

08
Lab Exercise
Lab Exercise

Stream structured tool_use output - Lab Exercises

1h 5m 1 Exercises

By the end of this hands-on lab you will have consumed a streaming Claude tool_use response — watching the tool_use.input JSON arrive field-by-field via content_block_delta events, using a streaming JSON parser to render each field into a live console UI as it lands, and measuring the first-field-time vs total-time win over blocking mode. Uses the PythonAI container + Claude API proxy. Proves streaming structured output feels dramatically faster for user-facing extraction UIs.

09
AI Lesson
AI Lesson

Extraction eval: precision / recall on structured fields

1h 0m

By the end of this lesson you will know how to score an extraction pipeline: per-field precision (of the fields we extracted, how many were right), per-field recall (of the fields present in ground truth, how many did we get), exact-record match rate (all fields correct for one record), and when to reach for LLM-as-judge fallback (semantic fields where string equality fails). Anchors the metrics you'll compute in the L10 capstone scorecard.

10
Lab Exercise
Lab Exercise

Extraction pipeline capstone - Lab Exercises

1h 15m 1 Exercises

By the end of this hands-on capstone you will have built the full CLD-AI-108 deliverable — a 10-ticket extraction pipeline using tool_use with the invoice-style nested schema from L4, batched with concurrency + retry from L6, and scored against gold labels for per-field precision + exact-record match rate. Emit a JSON scorecard defensible in a code review. Uses the PythonAI container + Claude API proxy — pulls together every technique from CLD-AI-108.

This course includes:

  • 24/7 AI Instructor Support
  • Live Lab Environments
  • 5 Hands-on Lessons
Skill Level Intermediate
Total Duration 10h 45m