From setting up Ollama, through tokens and temperature, the Chat API, prompt design, tool use, RAG, and production operations, all the way to a FastAPI vehicle-info assistant. Across 50 lessons you'll get to the point where you can build your own AI app on top of an off-the-shelf LLM. No paid API key required — every piece of code was actually run against a local LLM on your own machine and public automotive data.
The 50 lessons are split into 7 chapters. We recommend working through them in order starting from Chapter 1, but feel free to skim just the parts you're curious about. * Code that uses Ollama or the network can't run in the browser, so try it in the local environment (Ollama and Python) you set up in lesson 2. Only parts written with the standard library alone can run in the browser.
Start by running a local LLM with Ollama on your own machine, no paid API key required. Cover tokens, output length, temperature and seed, and how to pick a model — the fundamentals for treating an LLM as a component.
Cover the whole surface of the conversation API: the system, user, and assistant roles, multi-turn conversations, streaming, and JSON output. We finish by building a terminal chat that keeps its own history.
Learn the patterns for assembling a prompt as a component: assigning a role, few-shot examples, XML tags, and locking down the output format. We finish by building a small evaluation dataset and improving a prompt by comparing its quality with numbers.
Define a tool, receive a tool_call from the model, run it against the public automotive API (vPIC), and feed the result back through to a final answer. We finish by building a minimal agent with a loop that selects, runs, and answers.
Turn documents into embedding vectors, search them with cosine similarity, chunk long documents, and answer based on retrieved recall notices. Cover cited answers and evaluating retrieval, then the pitfall of not letting the model answer things the documents don't say.
Add every safeguard you need before going live, one at a time: cost estimation, rate limiting and retries, timeouts, logging, prompt injection and personal data, and output guardrails. We finish with a cache that takes advantage of deterministic output.
Combine everything you've built to design a vehicle-info assistant that routes questions by type, and expose it to the web as FastAPI endpoints. Wire VIN decoding and recall RAG into the endpoints, and close out by confirming the whole thing with tests.
Once you finish all 50 lessons, move on to Shipping and Running a Service, where you'll put this assistant on a server and publish it. All courses unlock with a membership.