Free through lesson 3

Intro to AI App Development

From setting up Ollama, through tokens and temperature, the Chat API, prompt design, tool use, RAG, and production operations, all the way to a FastAPI vehicle-info assistant. Across 50 lessons you'll get to the point where you can build your own AI app on top of an off-the-shelf LLM. No paid API key required — every piece of code was actually run against a local LLM on your own machine and public automotive data.

Curriculum

The 50 lessons are split into 7 chapters. We recommend working through them in order starting from Chapter 1, but feel free to skim just the parts you're curious about. * Code that uses Ollama or the network can't run in the browser, so try it in the local environment (Ollama and Python) you set up in lesson 2. Only parts written with the standard library alone can run in the browser.

Chapter 1 — What Is an LLM (lessons 1–6)

Start by running a local LLM with Ollama on your own machine, no paid API key required. Cover tokens, output length, temperature and seed, and how to pick a model — the fundamentals for treating an LLM as a component.

Chapter 2 — Mastering the Chat API (lessons 7–13)

Cover the whole surface of the conversation API: the system, user, and assistant roles, multi-turn conversations, streaming, and JSON output. We finish by building a terminal chat that keeps its own history.

Chapter 3 — Prompt Design (lessons 14–20)

Learn the patterns for assembling a prompt as a component: assigning a role, few-shot examples, XML tags, and locking down the output format. We finish by building a small evaluation dataset and improving a prompt by comparing its quality with numbers.

14

The Basics of a Good Prompt (Assigning a Role)

The start of Chapter 3. How to write with_role, which assigns a persona before asking for something, and the three foundations of a good prompt: role, context, and a concrete instruction.

🔒 Basic
15

Aligning Output with Examples (Few-Shot)

Few-shot prompting has the model imitate input-output examples when an instruction alone gives inconsistent output. Using fuel-type classification, confirm 2-3 examples are enough to align the format.

🔒 Basic
16

Separating Instructions from Data with XML Tags

A way to separate content so an "ignore previous instructions" buried in external data isn't followed. Build it with xml_tag and build_context, plus a first-pass detection of instruction overrides.

🔒 Basic
17

Locking Down Output Format with a JSON Schema

Lock down keys that format="json" alone can't constrain, using a JSON schema. Any key listed in required is guaranteed to be present, so downstream code can read it with confidence.

🔒 Basic
18

A Prompt for the Automotive Domain (Passing Specs as Reference)

Instead of letting the LLM guess, hand it facts to ground its answer. Pass Tesla Model 3 specs decoded by vPIC as reference material, until it answers the fuel type as "electric."

🔒 Basic
19

Building an Evaluation Dataset

The foundation for turning "it feels better" into a number. An evaluation dataset of input/expected-output pairs, accuracy, and how to verify the evaluator itself.

🔒 Basic
20

Evaluating and Improving a Prompt

The cycle of measuring a prompt against your evaluation data and fixing it. Measure a few-shot-assisted prediction with accuracy, and confirm a passing bar of at least 2 out of 3.

🔒 Basic

Chapter 4 — Tool Use (lessons 21–28)

Define a tool, receive a tool_call from the model, run it against the public automotive API (vPIC), and feed the result back through to a final answer. We finish by building a minimal agent with a loop that selects, runs, and answers.

21

What Is a Tool (Letting an LLM Use Outside Tools)

The start of Chapter 4. The idea behind tool use: letting an LLM, which only knows what it learned during training, ground its answers using an outside tool — vPIC, this course's tool.

🔒 Basic
22

Defining a Tool (Name, Description, Argument Schema)

How to describe a tool to the model in JSON. The roles of name, description, and parameters, and how the description you write determines how accurately the tool gets selected.

🔒 Basic
23

Receiving a tool_call

Pass a tool with a question, and the model, instead of answering itself, replies "use this tool with these arguments." What's inside tool_calls, and how to check just the name and arguments.

🔒 Basic
24

Executing the vPIC Tool

Turn the model's nomination into real data. Look up a function by name with run_tool, and decode a VIN or fetch models via NHTSA's vPIC (no key required). Keep tests deterministic with a fixture.

🔒 Basic
25

Closing the Loop by Returning the Result

Return the tool's result to the model as a role: tool message and let it write the final answer. How to read run_agent, which closes the loop of nominate → run → answer.

🔒 Basic
26

Multiple Tools and Selection

Let the model pick the right tool for the question when there are two or more available. The description that guides the choice, and how to check what got used via steps.

🔒 Basic
27

Tool Errors and Retries

Handle tool failures as two kinds. Broken input, like a VIN that doesn't exist, is caught and returned to the model; transient failures get absorbed with an exponential-backoff retry.

🔒 Basic
28

The Minimal Form of an Agent

The minimal form of an agent that repeats selecting a tool, running it, and reconsidering based on the result. The loop inside run_agent, and how max_steps stops it from running away.

🔒 Basic

Chapter 5 — RAG (Retrieval-Augmented Generation) (lessons 29–36)

Turn documents into embedding vectors, search them with cosine similarity, chunk long documents, and answer based on retrieved recall notices. Cover cited answers and evaluating retrieval, then the pitfall of not letting the model answer things the documents don't say.

29

What Is RAG (Grounding Answers in External Documents)

The start of Chapter 5. The big picture of RAG: search for documents the model doesn't know, bring them into the prompt, and have it answer based on them. The material here is 5 recall notices.

🔒 Basic
30

Embeddings

The heart of RAG: embeddings turn a sentence into a meaning vector, via the embed function calling /api/embed. Confirm nomic-embed-text produces deterministic 768-dimensional vectors.

🔒 Basic
31

Searching with Cosine Similarity

Measure vector closeness with cosine similarity — 1 for the same direction, 0 for orthogonal. A search is just sorting each document's score against the question, highest first.

🔒 Basic
32

Chunking

Cut a long document to a searchable size. Write chunk to split by paragraph and further slice overly long paragraphs by character count, and see how granularity affects search accuracy.

🔒 Basic
33

Retrieving Recall Notices with RAG

Confirm that an index of recall notices retrieves the right document for each question. The strength of embedding search: it retrieves by meaning even when the wording differs from the document.

🔒 Basic
34

Answering with Citations

RAG's finishing touch: answer_with_citation grounds answers only in retrieved material and cites source IDs in brackets, plus an instruction against inventing anything not in the documents.

🔒 Basic
35

Evaluating Retrieval (Hit Rate)

When a RAG answer is bad, is the cause retrieval or generation? Measure retrieval on its own with hit rate — how often the expected document lands in the top k — and fix retrieval before generation.

🔒 Basic
36

A RAG Pitfall (Don't Let It Pick Up What Isn't in the Documents)

RAG errs when it forces a retrieval and answers from it anyway. Use the fact that even an unrelated question's top result scores low, plus a threshold to decide when to say "not in the documents."

🔒 Basic

Chapter 6 — Production Operations (lessons 37–44)

Add every safeguard you need before going live, one at a time: cost estimation, rate limiting and retries, timeouts, logging, prompt injection and personal data, and output guardrails. We finish with a cache that takes advantage of deterministic output.

37

Calculating Tokens and Cost

Chapter 6's start. A local LLM costs nothing in API fees, but estimate cost anyway in case you move to the cloud. How to calculate token billing, where input and output have different unit prices.

🔒 Basic
38

Rate Limiting and Retries

Contain runaway cost from abuse or bugs with rate limiting that caps requests in a time window. Absorb transient failures with backoff retries, and raise an exception to the caller when those fail.

🔒 Basic
39

Timeouts and Concurrency

Give every call a timeout so requests don't pile up waiting on a slow response. The idea of capping concurrency to protect both your server and your costs.

🔒 Basic
40

Logging and Observability

In production, record what you used, how much, and how fast — but never leave personal data in the log. The shape of make_log, which routes everything through redact_pii before recording it.

🔒 Basic
41

Defending Against Prompt Injection

Preparing for an attack that hides "ignore previous instructions" in external data. First-pass detection with looks_like_injection, isolation via XML tags, and covering what those still miss.

🔒 Basic
42

PII and Safety (Masking Email and Phone Numbers)

Mask personal data hiding in a user's input before it ever reaches the LLM or the log. redact_pii's regular expressions, and what they do and don't catch.

🔒 Basic
43

Validating Output (Guardrails)

Even with a schema, an LLM's output isn't perfect. A guardrail that checks required keys with validate_json_shape before passing it downstream, retrying or erroring out if any are missing.

🔒 Basic
44

Caching and Reproducibility

Because temperature=0 and a fixed seed make output deterministic, it's safe to cache and reuse an answer for the same input. The Cache class, and confirming the LLM never runs on a second hit.

🔒 Basic

Chapter 7 — Finishing the App (lessons 45–50)

Combine everything you've built to design a vehicle-info assistant that routes questions by type, and expose it to the web as FastAPI endpoints. Wire VIN decoding and recall RAG into the endpoints, and close out by confirming the whole thing with tests.

Once you finish all 50 lessons, move on to Shipping and Running a Service, where you'll put this assistant on a server and publish it. All courses unlock with a membership.