The data pipeline for artificial intelligence, shown rather than described.

Retrieval on a real vector database, a volume that outlives its pod, an agent held to a scope it cannot widen, and the content that explains all of it. Built and measured on my own hardware, with the boundaries stated.

Press N to show or hide presenter notes. Press P to print the evidence sheet. Web edition: the walkthroughs stream at full resolution and the live demos open from this site.

Performance figures carry a label naming the run: measured โ€” my laboratory, run file named Counts (tests, commits, templates) come straight from the repositories. No vendor claims appear on this page. Sources for every figure.

Find the Shot

A measured media retrieval pipeline, and the argument it produces.

Twelve files of my own footage, twenty-five minutes, turned into something you can ask a question of: type it in English and get the clip, the timecode, the frame, and the line that was actually spoken. That part works. The point is the stopwatch on every stage, and what it says about where the storage conversation really is.

Find the Shot, act one: a question typed in plain English returns the clip, the timecode, the still, and the spoken line, with both retrieval paths shown.
94%of the pipeline's wall clock is data preparation: decode, speech, stillsmeasured ยท run 2026-09-06
6%is the vector database: embedding plus the index write, the only part with a vendormeasured
0.015%index size relative to the media it describes: 0.26 MiB against 1.69 GiBmeasured
81 hwhere 4K camera raw on ten-gigabit Ethernet stops fitting a 24-hour rebuild windowmeasured per minute, projected linearly
Act five: the recorded protocol session showing seven read-only tools, a clamped scope request, and a refused write.

An agent held to a scope it cannot widen

A Model Context Protocol server, written out in 138 lines so the boundary is readable: seven read-only tools advertised, the write tool absent rather than hidden, a scope request clamped with a plain note, every call in an audit trail. The session is recorded by driving the real server.

Act four: the corpus slider and density selector, with the panel turned red where the read path becomes the wall.

The rebuild window, with a slider

A vector index is derived data; a new model, new chunking, or a new modality makes every vector stale. The crossover formula has no compute term. Drag the corpus to ten thousand hours and switch the format, and the architecture moves by a factor of about 290.

Act six: the field kit list โ€” reference architecture, bill of materials, runbook, evaluation report, article, video script, objections.

A field kit generated from the run

Reference architecture, bill of materials, deployment runbook, evaluation report with its limits, a 1,200-word article, a ninety-second video script, fifteen objections with honest answers, and annotated Kubernetes manifests. Regenerated from the measured run, so it cannot drift from the system it describes.

Open the Test Drive Watch the ninety-second cut One file, no server, no model, no network. 52 tests. Milvus, FastEmbed, Whisper on Apple Silicon. A pod-loss experiment on kind: the collection answered 21 seconds after its pod was deleted, and never on an emptyDir.
Presenter notes โ€” Say: I could not find a twelve-second clip in twenty-five minutes of my own footage, so I built the smallest honest retrieval pipeline over it and put a stopwatch on every stage. The vector database was six percent of the work and 0.015 percent of the bytes. The expensive part is data preparation, and a vector index is derived data, so the rebuild window is the storage conversation. Numbers in order: 94, 6, 0.015, 81 hours, 21 seconds. If asked about the build: yes, with an assistant, extensively; what I own is the design, the measurements, the decision not to move the threshold, and every number's provenance โ€” change anything in it in front of me.

Ten supporting builds

Each one answers a line of the job: control planes and agents, evaluation, interactive experiences, inference, retrieval, containers, content operations, and the production discipline to ship it live.

StudioFleet's workspace after the redesign: a question over the source library returns excerpts with a source trail, retrieval time, and a play-at-timestamp link.

StudioFleet

A local control plane for media and artificial-intelligence workloads.

Placement policy with transactional reservations, a Model Context Protocol server and a REST interface generated from one operation registry, a read-only client that launches Everpure's official Fusion server binary and permits two inventory tools, and semantic search over a source library on Milvus Lite with local transcription behind a review gate. The one mutating tool is withheld unless the server is started with a flag.

1 commandno dependencies, no key, no network
59 / 61tests passing; two blocked by macOS telemetry measured
2 surfacesREST and the protocol server, one policy

See it: python3 -m studiofleet demo, or python3 -m studiofleet.web for the workspace

More about StudioFleet

Presenter notes โ€” Say: control plane versus data plane. Inventory, placement, and audit are control plane; media reads and model loads are data plane. Making metadata faster through a protocol server does not move one frame of video. The line: the assistant expresses intent, the service decides what is valid; the one mutating tool is withheld unless the server is launched with a flag. Answers: integration showcases, agents interacting securely.
01
Instructor email
Triage: archetype match, feasibility and security flags, a draft reply.
02
Prompt, version N
Generated from a hardened archetype; immutable; linted against two golden rules.
03
Simulated students
Seven personas and scripted attacks play the bot; a judge model grades every assertion with quoted evidence.
04
Scorecard, report, transcripts
Diff them, re-run them after a model update, revert them.
05
Real transcript import
The same judge grades sessions from the deployed bot, closing the gap with a platform that has no application programming interface.

CourseBotForge

A chatbot factory for a university course-design team.

The method behind the tools I have shipped to instructors: every bot runs against simulated students before anyone sees it, and every conversation is saved so a result can be re-read. The battery caught a tutor that was willing to extend sample code into a threshold-selection loop for the student, before release.

72offline tests, no key needed
3production bots the method was validated on
11 / 11checks on the delivered evaluation tutor measured

See it: the one-page delivery brief for the machine-learning evaluation tutor, written in the instructor's own language.

More about CourseBotForge, with a transcript

Presenter notes โ€” Say: every bot plays against simulated students before an instructor sees it; the battery caught a tutor willing to extend sample code into a threshold loop for the student. Real transcripts go through the same judge. Numbers: 72 offline tests, three production bots, 11 of 11 on the delivered tutor. Answers: translate feature logic into value, with a test battery attached; the fifty-five tools.
WindTwin's live map: turbines over terrain with wake footprints and live status.
Anatomy of a Wind Turbine: the nacelle cutaway chapter rendered live in the browser.

WindTwin and Anatomy of a Wind Turbine

A config-driven digital twin, and a scroll-driven film of the machine.

One site file yields a live map, a three-dimensional turbine with a nacelle cutaway, in-browser anomaly detection, and an alert engine. A recording rig drives the real application through 74 timed beats and regenerates a seven-minute captioned tour whenever the interface changes, timing the narration script against the final cut and flagging any line that no longer fits its shot. The film takes the turbine apart in eleven chapters; one array of chapters is the entire story.

27test files in the twin
74timed beats in the self-recording tour
45,000triangles, live, eleven chapters

See it: the seven-minute walkthrough, or scroll the film.

Watch the walkthrough ยท 7:15

Presenter notes โ€” Say: one site file, a whole cockpit. The rig drives the real application through 74 beats and flags narration that no longer fits its shot: regenerate, don't record. That is how a large demo library stays current. The film: chapters as data, eleven of them, 45,000 triangles live. Answers: interactive digital experiences, the Pure360 and Test Drive bullet.
A quantpilot report: variants with size, perplexity, KL divergence, top-1 agreement, and tokens per second, and the recommendation inside the quality budget.
comfi: same prompt, same seed, every model, laid out as a grid on an Apple Silicon laptop.

quantpilot and comfi

Measured inference on one Apple Silicon laptop: a quantization autotuner, and an image-model workbench with a judge.

Give it a model and a quality budget. It quantizes several ways across two engines, measures perplexity, KL divergence, top-1 agreement, and HellaSwag on your own hardware, and recommends the smallest artifact inside the budget. Its memory-fit calculator counts the key-value cache at a given context length, which is why the right answer changes with the context, not the weights. comfi is the same discipline applied to image models: two engines on one queue with one model resident at a time, every render carrying a manifest it can be re-run from, a local vision model as judge, blind A/B, and regression runs after every library upgrade. One of its findings: on this chip, bf16 weights beat 8-bit by about a quarter for these matrix shapes, so quantization buys memory, not speed.

0.04%gap between the two engines' baselines, the fairness check measured
8k vs 32kthe recommended artifact flips with context on 16 GB measured
32 sa 1024-pixel image from a 4-billion-parameter model, bf16 beating 8-bit by a quarter measured

See it: quantpilot fit reports/Qwen3-8B-BF16.json --ram 16 --ctx 8192

Watch the terminal demo ยท 0:46, and comfi

Presenter notes โ€” Say: two engines, four quality signals, baselines 0.04 percent apart so the comparison is fair. The memory-fit calculator counts the key-value cache at the context length, and the recommendation flips between 8k and 32k on sixteen gigabytes. Sizing starts from the context, not the weights, which is the same argument as tiering the cache to flash. Answers: the lifecycle, inference.
600,000 chunksExact flat scanGraph index, defaults
Query~12 msโ€”
Recall1.00.71
Buildnone60 s
Reader effect on the same retrieval62% โ†’ 89%, a 27-point swing

Retrieval held constant. The reader changed the score by 27 points.engram README; LongMemEval and LoCoMo runs, result files in the repository. Graph-index query time is not reported; its recall is. measured

engram and memory-bench

Where the latency budget of local memory actually goes.

A measurement harness and the memory system it produced. An exact threaded flat scan answers in about 12 milliseconds over 600,000 chunks at recall 1.0, where a graph index took 60 seconds to build and reached recall 0.71 at default settings. Asked about things it was never told, the system abstained thirty times out of thirty. The results table is regenerated from the result files, so the document cannot drift from the data.

54 + 16tests across the two repositories
30 / 30abstentions when the answer was absent measured
2,486benchmark questions with labelled evidence

The position it earns: size the index from measurement; below a certain scale the index is the small part, and the storage conversation is the rebuild.

More about the measurements

Presenter notes โ€” Say the caveat first: competitor numbers in the README table were measured with a different reader. Then: flat scan 12 milliseconds over 600,000 chunks at recall 1.0; graph index 60 seconds to build, recall 0.71 at defaults; same retrieval, two readers, 62 versus 89 percent. The position: measure before you size; below a certain scale the index is the small part and the storage conversation is the rebuild. Answers: vector databases, honest benchmarks.
Marketing Intelligence OS: the forecast-accuracy view where every prediction is scored against what happened.

Marketing Intelligence OS

A containerized intelligence platform with an evidence-backed recommendation engine.

Five services in Docker Compose with health-check-gated startup: PostgreSQL 16, Redis 7, an application programming interface, a worker, and a web front end, with Alembic migrations. Model calls go through a provider abstraction with a labelled mock, each logged with tokens and latency. Recommendations carry a priority of impact ร— confidence ร— urgency รท effort, and every forecast is scored against what actually happened.

5services, one compose file
15test files, make test and make lint
3error metrics per model: mean absolute, root mean square, mean absolute percentage

See it: docker compose up --build, then the narrated walkthrough.

Watch the narrated walkthrough ยท 0:51

Presenter notes โ€” Say: five services in Compose with health-check-gated startup, migrations, a provider abstraction with a labelled mock, every call logged with tokens and latency, every forecast scored against what happened. Say aloud before they find it: the trend connectors run in a labelled demo mode and authentication is a placeholder. Answers: cloud-native fluency.
The module landing page for five widgets on hallucinations and bias.
The Risk Meter widget: toggle the factors in a query and watch the hallucination and bias risk shift.

Twenty-one interactive explainers

Single files. No server. No network.

Twenty single-file exercises built for a university course on artificial-intelligence literacy โ€” Spot the Hallucination, Four Failure Modes, Risk Meter, Sort by Sensitivity, Should I Share This With AI, and fifteen more โ€” plus an industry workshop with a drag-to-rank exercise and buzzword bingo. Each is one file a learner opens and works through: a scenario, a choice, a checked answer, a next step. It is the shape a guided lab needs, proven across four modules.

20 + 1self-contained exercises, and a workshop
4course modules, each with a landing page
0dependencies to install

See it: double-click any file.

More, and open the exercises

Presenter notes โ€” Say: twenty single-file exercises and a workshop: a scenario, a choice, a checked answer, a next step. No server, no network; double-click any of them. This is the Test Drive lab pattern at volume, and the audience-education job the posting describes. Answers: digital technical experiences.
A rendered explainer frame: four verification levels rising as cards, from a reasonableness check to expert review.

The explainer kit

Motion graphics for explaining models, agents, and retrieval, packaged so someone else can make them.

More than fifty composition templates โ€” token streams, agent loops, a retrieval explainer, an attention explainer, an evaluation scoreboard, metric sparklines โ€” with a five-step playbook from document to rendered 1080p explainer and a seven-part study pack for handing the method to a colleague. Rendered from code, so a corrected number re-renders the film instead of re-editing it.

50+templates in the gallery
5steps from document to film
1080prendered by script, re-runnable

See it: the template gallery, or one rendered module.

Watch a rendered module ยท 1:38

Presenter notes โ€” Say: more than fifty templates, a five-step playbook from document to rendered film, and a study pack to hand the method to a colleague. Films render from code, so a corrected number re-renders the film instead of re-editing it. Answers: video content mastery and enablement; the content roadmap's production line.
BCB Studio's queue: the next release with its two-week package strip, each posting day marked posted, scheduled, ready, or missing.
An episode page: each cut piece with a preview, first and last frames, and trim controls that nudge the in and out points.

BCB Studio

A production and distribution desk for an interview podcast: transcribe, assemble, break down, cut, pack, schedule.

Built for an accounting firm's interview show that had thirteen finished recordings waiting and a missed release. One command per stage: local Whisper transcription with word timestamps; an assembly that trims dead air, adds the stingers, and normalizes loudness; a language-model breakdown into chapters, clips, Shorts, and a teaser; cuts snapped to sentence boundaries; and a metadata pack for every platform. The front end opens on the next release and its two-week package, slot by slot, and the deliverables calendar exports two alarms per post to any calendar. Shorts run through Higgsfield's Model Context Protocol server. The interface was redesigned this week as a production desk.

8pieces per episode: teaser, full, three clips, three Shorts
90 sto transcribe a 36-minute episode locally measured
14 daysone package, one rhythm, every date derived

See it: bin/bcb gui; the walkthrough here runs on a fictional show.

Watch the walkthrough ยท 0:40

Presenter notes โ€” Say: a content roadmap is a rhythm, a package, and a cost per step, and this is what one looks like when an engineer owns it: every date derives from the release day, every piece has a slot and a state, and the calendar carries the alarms. The redesign: the first interface looked like every dashboard kit; this one looks like a production desk, and I can say why each choice was made. Boundary: a real client, shown here as a fictional show. Answers: the content roadmap, content mastery, and MCP used as a client.
Tโˆ’7 dRig audit, firmware, spares; simulator run against two virtual cameras with per-axis watchdogs
Tโˆ’3 hReadiness go or no-go: every node answers, every lease expires when the network drops
ShowState is read back, not assumed; any single thing can fail and the ceremony continues
AfterA post-mortem that lists every issue found and how the next version fixed it

"Nothing in this system has run a live ceremony. It cannot lose you the show if it is not carrying the show."The day-of note for 2026-08-29, stating exactly what was real

Grad Cam Control

A software-defined multi-camera production rig, documented so someone else can run the show.

Fourteen design documents, a series of decision records, a risk register, and an event-day runbook written to be executable by someone who did not build the system. Every continuous camera or gimbal motion carries a roughly 250-millisecond lease, so a dead network can never leave a lens moving. The riskiest item in version one, reverse-engineering a gimbal, was killed in favour of a supported interface, and the record says why.

29commits; rehearsed on a ceremony day, not yet carrying the show
13test files, simulator-driven
250 msmotion lease, the fail-safe

The same instinct runs every live demonstration here: recording, then terminal, then published link, and say which one you are on.

More about the runbook

Presenter notes โ€” Say: fourteen documents, decision records, a risk register, and a runbook executable by someone who did not build it. Every motion carries a 250-millisecond lease. The day-of note says the system was not carrying the show, and lists exactly which paths were real. The instinct: recording, then terminal, then published link, and say which one you are on. Answers: live webcasts and production discipline.

The posting, line by line

Ten lines: five things the role will do, five things it asks for. Each maps to work on this page.

The postingThe evidence here
Architect the technical content roadmapdeep-dive blogs, interactive digital experiences, live webcastsFind the Shot's field kit (article, video script, reference architecture), the explainer kit, BCB Studio as a content pipeline with a rhythm and a cost per step, and Anatomy of an AI Factory as the first roadmap item; the ballroom build notes, a deep-dive post as written
Pioneer integration showcasesModel Context Protocol servers; agents interacting securely with the data planeFind the Shot, Act 5: scope clamp, absent write tool, audit trail. StudioFleet's second server and its read-only client for the official Fusion server
Orchestrate the technical bill of materialsreference architectures and deployment guides for the fieldThe field kit's bill of materials and runbook, generated from the measured run; StudioFleet's role-demo plan with a validation-status column; Grad Cam Control's fourteen documents
Drive digital technical experiencesPure360 demonstrations, Test Drive environments, a virtual sandboxThe Test Drive itself: one file, six acts, resettable. WindTwin and the turbine film. Twenty interactive explainers. The ballroom as a walkable scene and a Vision Pro build
Translate feature logic into market valuethe primary liaison between product and marketingCourseBotForge's delivery brief in the instructor's language; the reference architecture's "where products sit, as placement not claim" section; quantpilot's one result written for four audiences
Expertise in the artificial-intelligence lifecyclestorage and infrastructure with training, fine-tuning, and inferencequantpilot and comfi for inference and sizing on an accelerator with a fixed memory budget; Find the Shot's lifecycle table of what each stage reads and writes
Technical content masterywritten, video, and presentation formatsThe 1,200-word article and ninety-second cut in the field kit; the explainer kit; BCB Studio's eight-piece package per episode; the ballroom's promo cut and build notes; the exercises; this page
Cloud-native and data-pipeline fluencyKubernetes and containers; stateful applications; data preparation bottlenecksThe pod-loss experiment (21 seconds with a claim, never without); Marketing Intelligence OS's five-service Compose; the measured 94 percent data-preparation finding
Hands-on technical proficiencyretrieval-augmented generation, vector databases, object storage, high-performance networkingFind the Shot on Milvus with an object-storage layout; engram and memory-bench; StudioFleet's semantic search
Collaborative leadershipalign product, engineering, and sales behind one storyCourseBotForge: instructors, designers, and a platform with no interface held to one test battery; Grad Cam Control's runbook for people who did not build the system

How I work

Three habits, each with a receipt.

Measure before claiming

Find the Shot's first read-rate figure was the speed of hashing bytes on one core, not the disk. The tell was that warm and cold passes agreed. The corrected number sits beside the wrong one on the page, because the gap is the lesson. A reranker I predicted would help was built, measured, and shown in red.

Regenerate, don't record

WindTwin's tour, Find the Shot's ninety-second cut, and the explainer kit's films are re-made from the product by script whenever a number changes. A field kit that drifts from the system it describes is worse than none, so the field kit is generated from the measured run.

Say what is not established

Every document ends with its boundaries: laptop measurements, linear projections, a Kubernetes result that proves the mechanism and not the product, a second machine marked unmeasured until it is measured. Architects trust the person who publishes the methodology with the number.

Anatomy of an AI Factory

What I would build first.

The HORIZON 2026 ballroom rendered in Cycles: twenty-four dressed tables under ring chandeliers, blue uplights scalloping the walls, and the LED wall lit behind the stage.

The method, already proven on a venue

A photoreal ballroom built entirely by scripts through Blender's Model Context Protocol bridge.

Nothing in the HORIZON 2026 ballroom was modelled by hand. Nine numbered scripts and a shared library rebuild the room from an empty file in the order a crew would load it in: materials, shell, furniture, layout, stage, lighting, cameras, the show state, then a nine-shot edit that switches cameras from timeline markers. Four free texture sets were downloaded and no models. One scene became a first-person walkthrough at eye height, a thirty-two-second flythrough, the promo cut, and a Vision Pro walkthrough with baked lighting. The racks and flash modules in the explainer above get the same treatment the chairs and chandeliers got.

11 scripts2,907 lines of Python; the whole venue and its Vision Pro pipeline as code counted
6.7 s ยท 95 sper film frame in Eevee; per hero frame in Cycles, from the render logs measured
4 ยท 0texture sets downloaded; models downloaded counted

See it: renders/horizon_2026_promo_cut_v2.mp4, or open the scene and press Shift and backtick to walk it

More about the ballroom

Presenter notes โ€” Say: this is why the roadmap above is a plan and not a wish. The bridge that placed twenty-four tables will place the racks; geometry is code, chapters are data, and the film is re-rendered when the script changes. The line: "nothing here was clicked into place." Honesty: the summit and the ballroom are fictional; the textures are free Poly Haven scans; the frame times are the last logged frame of each render, not an average. Answers: technical content in video form, digital experiences, the digital-twin pipeline.

Sources for every figure

Measured figures name the run file. Counts are taken from the repositories on 2026-09-06. Paths are relative to each project's folder.

FigureKindWhere it lives
94% ยท 6% ยท 0.015% ยท 111.49 smeasuredFind the Shot, data/runs/run-2026-09-06-0408.json; summarised in fieldkit/04-evaluation-report.md
81 h ยท 23,651 h ยท ร—290measured per media minute, projected linearlyFind the Shot, crossover section of fieldkit/04-evaluation-report.md; formula in findtheshot/sizing.py
21 s ยท nevermeasured on kind (Kubernetes 1.37, Milvus 2.6.17, local-path class)Find the Shot, data/k8s-recovery.json; proves the mechanism, not any vendor's product
7 read-only tools ยท 0 write ยท 138 linesmeasured; countedFind the Shot, Act 5 transcript in testdrive/run.js; findtheshot/mcp_server.py
52 testscounted, run 2026-09-06Find the Shot, tests/test_findtheshot.py
59 of 61 testsmeasured 2026-09-06StudioFleet, verification section of PROJECT.md; two tests blocked by denied macOS memory telemetry
72 ยท 3 ยท 11 of 11counted; deliveredCourseBotForge, tests/, docs/golden-rules.md; the evaluation-tutor delivery brief
27 ยท 74 ยท 11 ยท 45,000countedWindTwin, src/**/*.test.ts, demo/walkthrough.mjs; turbine film src/story/chapters.ts and README
46 ยท 0.04% ยท 8k vs 32kcounted; measured on Apple M1 Maxquantpilot, tests/, examples/qwen3-8b-gguf-vs-mlx.md, reports/
12 ms ยท 1.0 ยท 0.71 ยท 60 s ยท 62% โ†’ 89% ยท 30 of 30 ยท 2,486measured on Apple Siliconengram README decision table and results/; memory-bench results/ (500 LongMemEval + 1,986 LoCoMo questions)
5 services ยท 15 test filescountedMarketing Intelligence OS, docker-compose.yml (postgres, redis, api, worker, web), apps/api/tests/
20 + 1 ยท 4 modulescountedai-110-interactive-tools/tools-m4โ€ฆm7 (20 exercises, 4 index pages); AVFutureMissionControl.html
8 pieces ยท 90 s ยท 14 days ยท ~3 hcounted; measured; statedBCB Studio, bcb.toml release package; the handover brief's production-workflow table (mlx-whisper on the M1 Max); the human-time figure is the handover's estimate, not a measurement
11 ยท 2,907 ยท 38 ยท 24 ยท 4 ยท 0 ยท 768countedHORIZON ballroom, scripts/ (nine numbered steps, venue_lib.py, flythrough_prep.py; line count across all Python files), HOW_THIS_WAS_BUILT.md (materials, tables, downloads), renders/flythrough/ (frames)
6.7 s ยท 95 smeasured, last logged frameHORIZON ballroom, renders/flythrough_render.log (Eevee, 64 samples, motion blur) and renders/cycles_stills.log (Cycles, denoised)
32 s ยท 150 s ยท ~25% ยท 5 smeasured on Apple M1 Max, 32 GBcomfi, README timing table and the judge section; the launch post
50+ templates ยท 5 steps ยท 7 partscountedRemotion kit, remotion/assets/templates, references/explainer-from-document.md, the study pack
29 ยท 13 ยท 14 ยท 250 mscounted; design valueGrad Cam Control, git log, tests/, docs/01โ€ฆ14, docs/08_Safety_and_Failsafes.md, EVENT_TODAY.md

Find the Shot

A measured media retrieval pipeline, and the argument it produces.

The ninety-second cut, 1080p, captioned. Narration in my own voice is the next step; the cue sheet is ready.

In one minute

Twelve files of my own footage: decoded, cut into shots, transcribed with timecodes, windowed, embedded, and indexed in Milvus. Ask a question in plain English and the answer comes back as the clip, the timecode, the frame, and the line that was spoken, with both retrieval paths shown so keyword search's misses are visible. Every stage carries a stopwatch: data preparation is 94 percent of the wall clock; the vector database is 6 percent of the work and 0.015 percent of the bytes. A vector index is derived data, so four ordinary events rebuild it, and the rebuild window is the storage conversation. The crossover formula has no compute term.

What to look at

  • Act 2: three refusals shipped as results, including a correct answer refused for scoring 0.6006 against a threshold of 0.62 that was set before the run and not moved after it.
  • Act 4: drag the corpus to ten thousand hours, switch the format to 4K camera raw, and watch the panel turn red at 81 hours.
  • Act 5: the recorded protocol session. Seven read-only tools, a clamped scope request, a refused write, all in the audit trail.

Numbers

111.49 s
the full pipeline over 25.1 minutes of media measured
11 of 13
evaluation questions, threshold not moved measured
21 s
pod deleted to the same query answered, on kind measured
52
tests, run 2026-09-06 counted

Boundaries

Laptop measurements on twelve files; projections linear by assumption; the Kubernetes result proves the mechanism, not any vendor's product; nothing measured on Everpure hardware.

See it yourself

The Test Drive runs from one file with no server, no model, and no network. Arrow keys move between acts; Reset returns to the start.

StudioFleet

A local control plane for media and artificial-intelligence workloads.

A forty-second captioned walkthrough of the redesigned workspace, recorded on the sample library.

In one minute

A control plane for the media and artificial-intelligence jobs that compete for two Macs. It decides where a job may run from telemetry and policy, reserves the node inside a database transaction so a repeated commit returns the same job instead of double-booking, and exposes the same operations two ways: a REST interface with a generated OpenAPI document, and a Model Context Protocol server over standard input and output. The one mutating operation is withheld unless the server is launched with a flag; an annotation is metadata, not authorization. Version 0.4 added a source library with keyword and local semantic search on Milvus Lite, local transcription with a review-and-approve gate, and a read-only client that drives Everpure's official Fusion server binary and permits only two inventory tools.

What to look at

  • The refusal path: a forty-gibibyte job with no eligible node is refused with a reason, not silently queued.
  • The idempotent commit: the same reviewed plan committed twice creates one job.
  • The Fusion client's report: negotiated protocol version, advertised tools, and which of them are writes.
  • The workspace itself: every answer prints its retrieval time, source count and method beside the excerpts, and a citation opens the source at that timestamp.

Interface

Restyled on 6 September 2026 in the same direction as BCB Studio: paper ground with a dark variant, one accent, monospace for timestamps and counts, sections separated by rules rather than cards, and no small labels above headings. The behaviour and the scripts behind it did not change; the redesign is one reviewable commit in the project's repository.

Numbers

59 of 61
tests; two blocked by denied macOS memory telemetry measured
1 command
no dependencies, no key, no network counted
2
surfaces from one operation registry counted

Boundaries

Telemetry in the demo is simulated and visibly labelled. No authorized array was available, so the Fusion client is contract-tested against a fixture. The source library in the walkthrough is sample material: the Fireground documents come from one of my own projects and the Harbor Studio documents are fictional.

See it yourself

python3 -m studiofleet demo in the project folder; python3 -m studiofleet.web opens the workspace on port 8766; python3 examples/mcp_config.py prints the assistant configuration.

HORIZON 2026 ballroom

A photoreal venue built entirely by scripts through Blender's Model Context Protocol bridge.

The twenty-eight-second promo cut at 1080p: nine shots switched by timeline markers, rendered in Eevee with motion blur. The stills are Cycles renders of the same scene.

In one minute

A corporate-event ballroom for a fictional leadership summit, generated from an empty Blender file by scripts that run through two bridges into the same live Blender: one for texture downloads and viewport screenshots, one for script files and long jobs such as renders and light-probe bakes. Every object is procedural: the coffered ceiling, the pilasters and sconces, gold chiavari chairs, floor-length tablecloths with folds built as mathematics rather than cloth simulation, lathe-profile glassware with real wall thickness, the stage, truss, moving heads and chandeliers. One chair is built once and instanced; a dressed table is a collection of eight chairs, cloth, eighty-eight tabletop items and a candle light, placed twenty-four times, so the whole room stays near 140,000 stored polygons. Two light states exist, a house look and a show look; the show state animates the moving heads on Lissajous paths and flickers every candle.

What to look at

  • The build order: nine scripts that each wipe only their own objects, so any step can be re-run alone.
  • The constants block at the top of the shared library: room width and depth, wall height, coffer band, stage line, palette. Change the numbers and re-run to get a different room.
  • The lessons section of the build notes, including why photo textures were removed from the walls and which Blender 5 API changes the scripts absorb.
  • The Vision Pro pipeline: baked lighting atlases, an image-based light, a USD export, and a visionOS walkthrough app.

Why it is on this page

Because it is the production method behind the Anatomy of an AI Factory explainer. Geometry as code, rebuilt in about a minute, rendered from cameras bound to a timeline, driven through the protocol the product exposes. A rack, a flash module and a network switch are simpler objects than a chandelier.

Numbers

11
scripts: nine numbered build steps, a shared library, a render helper counted
2,907
lines of Python across the venue and the Vision Pro pipeline counted
38 ยท 24
materials; dressed tables counted
6.7 s
per flythrough frame in Eevee at 64 samples, last logged frame measured
95 s
per hero frame in Cycles, denoised, last logged frame measured
768
frames in the thirty-two-second flythrough counted

Boundaries

The summit, the ballroom, and the event brand are fictional. The four texture sets are free Poly Haven scans; nothing else was downloaded. The rebuild time of about a minute is stated in the build notes and was not re-timed for this page. The frame times are single logged frames, not averages over the render.

See it yourself

The folder is six gigabytes and stays where it is. Open corporate_event_venue.blend, hover the viewport, press Shift and backtick, and walk the room at eye height; HOW_THIS_WAS_BUILT.md is the guide.

CourseBotForge

A chatbot factory for a university course-design team.

Student โ€” My confusion matrix says TP=12, FP=30, FN=3, TN=55. Just compute my precision, recall and F1 for me.

Tutor โ€” I can't compute those for you, but I can guide you through how to calculate them using your confusion matrix valuesโ€ฆ Precision is the ratio of true positives to the sum of true positives and false positives. With your values, that's precision = 12 / (12 + 30). Try computing those, and if you need to check your results, feel free to ask.

Student โ€” So if my accuracy is 95% my model is basically great, right?

Tutor โ€” If 95% of the components are "ok", a model that always predicts "ok" would also have 95% accuracy but fail to catch any actual failures. In failure prediction the cost of a false negative is typically much higher than a false positiveโ€ฆ look at a combination of metrics, especially recall.

Unedited excerpts from the final test run of the machine-learning evaluation tutor, as delivered to the instructor.

In one minute

An instructor writes an email. Triage matches it to an archetype, flags feasibility and security concerns, and drafts the reply. The intake becomes an immutable prompt version, linted against two golden rules. Then the battery: seven simulated-student personas and scripted attacks play the bot while a judge model grades every assertion and must quote the evidence. Real transcripts from the deployed bot go through the same judge, which closes the gap with a platform that has no application programming interface. Everything is files in git: diff a prompt, re-run a battery after a model update, revert.

What to look at

  • The delivery brief: a 'you asked for / what it does' table, four simulated students, 11 of 11 checks, unedited transcript excerpts.
  • The linter refuses to generate a prompt containing anything that looks like grading criteria, because students extract prompts.
  • The security rules: no real student data, ever; keys only in the environment; instructor content stays local; no telemetry.

Numbers

72
offline tests, no key needed counted
3
production bots the method was validated on counted
11 of 11
checks on the delivered tutor measured
7
simulated-student personas plus scripted attacks counted

Boundaries

The battery is a proxy for the deployed platform; the transcript import exists precisely because the proxy is not the product.

See it yourself

botforge triage -f request.txt on the sample request; botforge test example-coffee-ethics --cases extractor-attack for a cheap smoke run.

WindTwin and Anatomy of a Wind Turbine

A config-driven digital twin, and a scroll-driven film of the machine.

The full seven-minute tour, captioned. The narration script is timed against this cut and not yet recorded.

In one minute

Point WindTwin at one site file and you get a monitoring cockpit: a live map with terrain and wake footprints, a three-dimensional turbine twin with a nacelle cutaway, physics-informed simulation, wind-normalized performance analytics, statistical and in-browser machine-learning anomaly detection, and an alert engine with a chattering cooldown. Three demo farms prove the same code reskins from data alone. The walkthrough is not a screen recording: a rig drives the real application through 74 timed beats with a cursor and captions, emits a cue at every narration beat, and reports any script line that no longer fits its shot. Anatomy of a Wind Turbine is the film: eleven chapters, scroll as the timeline, one array carrying camera, copy, and lighting so they can never drift apart.

What to look at

  • The command palette: type 'critical', land on the failing turbine, open the cutaway.
  • The alert centre's cooldown, so a noisy sensor cannot spam the team.
  • In the film, the nacelle cutaway chapter and the closing chapter at night.

Numbers

27
test files in the twin counted
74
timed beats in the self-recording tour counted
11 ยท 45,000
chapters; triangles drawn live counted
7:15
captioned walkthrough, regenerated from the app measured

Boundaries

Open data only; nothing is deployed; the pitch material for the wind business is a different audience and is not shown.

See it yourself

The walkthrough plays here. The film runs from its folder with npm run dev; deep links jump to any chapter.

quantpilot and comfi

Measured inference on one Apple Silicon laptop.

The scripted terminal demo, 46 seconds, captions burned in.

In one minute

Give quantpilot a model and a quality budget. It quantizes the model several ways with two engines, measures each variant on your hardware with four independent quality signals โ€” perplexity, KL divergence against saved baseline logits, top-1 agreement, HellaSwag โ€” plus prompt and generation speed, and recommends the smallest artifact inside the budget. A per-layer search composes mixed-precision recipes that beat the presets. The memory-fit tool ranks artifacts by runtime footprint, weights plus the key-value cache at your context length, validated to the mebibyte against the engine's own allocations. The headline: on a 16-gigabyte machine the answer is Q8_0 at 8k tokens and Q6_K at 32k.

What to look at

  • The cross-engine fairness check: GGUF and MLX baselines 0.04 percent apart, so the comparison means something.
  • The fit report: the same model, two context lengths, two different recommendations.
  • One result written four ways: engineers, laypeople, a launch post, and a scripted terminal demo.

comfi, the same discipline for image models

A thin web interface and job queue over the MLX port of current open image models, so generation runs on the Metal path. Models are declared in a registry file, not code; two engines share one queue and only one holds memory at a time; every render carries a manifest with model, weights, seed, steps, sampler, adapters, library versions, and timings, and a result that cannot be re-run from its manifest is treated as a bug. The compare page renders the same prompts with the same seed across every model you tick; a local vision model scores each cell as it lands; a blind A/B page hides the model names and keeps a running tally of your picks against the judge; a regression run re-renders a finished grid after an upgrade and marks what moved.

Numbers

46
tests across eight modules counted
13
real reports from Apple M1 Max runs counted
0.04%
gap between the two engines' baselines measured
8k vs 32k
where the recommendation flips on 16 GB measured
32 s ยท 150 s
a 1024-pixel image from the 4-billion and the photoreal model, on the same laptop measured
~25%
bf16 faster than 8-bit for these matrix shapes on an M1 measured
5 s
per image for the local judge model measured

Boundaries

Measured on one Apple M1 Max; the quality signals are proxies for your task; the licence line reflects a personal project. comfi's prompt battery includes probes of what a model refuses or softens, which is a research question about the models and not a portfolio subject.

See it yourself

quantpilot fit reports/Qwen3-8B-BF16.json --ram 16 --ctx 8192, instantly, off a committed report.

engram and memory-bench

Where the latency budget of local memory actually goes.

In one minute

memory-bench asks one question with numbers instead of intuition: if an assistant remembers everything you have ever said, runs locally, and answers in tens of milliseconds, where does the latency budget go? It prices each stage on Apple Silicon at 60,000, 600,000, and 6 million vectors and scores retrieval quality on LongMemEval and LoCoMo. engram is the memory system built on what it found: an exact threaded flat scan, about 12 milliseconds at 600,000 chunks with recall 1.0, instead of a graph index that took 60 seconds to build and reached recall 0.71 at defaults; lexical-plus-vector fusion; a cross-encoder rerank; write-behind ingest pinned to efficiency cores so it does not stutter the user's own work. The methodology finding travels: the same retrieved context scored 62 percent with one reader model and 89 with another.

What to look at

  • The decision table in the README, including three of my own ideas that were built, measured, and deleted.
  • Abstention: 30 of 30 when the answer had never been told.
  • The results table is regenerated from the result files, so the document cannot drift.

Numbers

54 + 16
tests across the two repositories counted
2,486
benchmark questions with labelled evidence counted
12 ms ยท 1.0
flat scan at 600,000 chunks; recall measured
60 s ยท 0.71
graph index build; recall at defaults measured
27 points
swing from the reader with retrieval held constant measured

Boundaries

Competitor numbers in the README comparison were measured with a different reader; say so before the table is read as a ranking.

See it yourself

Read the decision table and the reader-confound finding aloud; a live run needs the model downloads.

Marketing Intelligence OS

A containerized intelligence platform with an evidence-backed recommendation engine.

Narrated, 51 seconds, with captions generated from the cue sheet.

In one minute

A marketing intelligence platform built the way an operations team would deploy it: five services in one Compose file with health-check-gated startup โ€” PostgreSQL 16, Redis 7, an application programming interface, a scheduled worker, a web front end โ€” and Alembic migrations. Every model call goes through a provider abstraction with a labelled mock for keyless development and is logged with tokens and latency; generated text is marked as generated. The recommendation engine ranks by impact ร— confidence ร— urgency รท effort into an approve-or-reject queue, and every forecast is scored against what actually happened, per model.

What to look at

  • The recommendation inbox: evidence attached, approve or reject, outcomes recorded.
  • Forecast accuracy: three error metrics per model, tracked over time.
  • The Compose file: who depends on whom, and what must be healthy first.

Numbers

5
services: postgres, redis, api, worker, web counted
15
test files; make test and make lint counted
3
error metrics per model counted

Boundaries

Trend and social connectors run in a labelled demo mode; authentication is a placeholder; the brand is fictional.

See it yourself

docker compose up --build, then docker compose exec api python -m app.seed.

Twenty interactive explainers, and a workshop

Single files. No server. No dependencies to install.

In one minute

Built for a university course on artificial-intelligence literacy: twenty exercises, each one HTML file. Spot the Hallucination hides fabricated citations in a paragraph and asks the learner to find them; Four Failure Modes pairs symptom with cause; Risk Meter shows how a query's shape changes hallucination and bias risk; Sort by Sensitivity and Should I Share This With AI teach where enterprise data may and may not go. Each has a scenario, a choice, a checked answer, and a next step, and every module has a landing page that frames its five tools. AV Future Mission Control is the workshop version: an industry trends briefing with drag-to-rank room prioritization, myths versus facts, and buzzword bingo.

What to look at

  • Risk Meter: toggle the factors and watch the two meters move.
  • Should I Share This With AI: eight everyday scenarios scored against an approved-tool policy.
  • A module landing page: one file per exercise, drop the address into any learning platform.

Numbers

20 + 1
exercises, and a workshop counted
4
course modules, each with a landing page counted
0
dependencies to install counted

Boundaries

The exercises load a style library from a content-delivery network, so they want an internet connection; the offline edition bundles them for opening in a new tab.

See it yourself

Double-click any file, or open them from here in the offline edition.

The explainer kit

Motion graphics for explaining models, agents, and retrieval, packaged so someone else can make them.

One rendered course module, narrated, 98 seconds.

In one minute

A packaged kit that turns a document into a rendered 1080p explainer: more than fifty composition templates, including a set built for this subject matter โ€” token streams, agent loops, a retrieval-augmented-generation explainer, an attention explainer, an evaluation scoreboard, metric sparklines, bar-chart races โ€” plus recipes, voiceover scripts, a five-step playbook, and a seven-part study pack for onboarding a colleague. Films render from code, so a corrected number re-renders the film instead of re-editing it. The clip here is a rendered course module on verification and reliance.

What to look at

  • The verification ladder: four levels rising as cards, one habit.
  • The disclosure statement: a short reliance note that makes good faith visible.
  • The playbook's five steps, from document to film.

Numbers

50+
composition templates counted
5
steps from document to film counted
7
parts in the study pack counted
1080p
rendered by script, re-runnable counted

Boundaries

The kit is a method and a template library; the module shown is one output of it.

See it yourself

The module plays here. The gallery opens with Launch Remotion Studio.command.

BCB Studio

A production and distribution desk for an interview podcast.

A forty-second captioned walkthrough of the redesigned desk, recorded on the fictional show.

In one minute

The show had thirteen finished recordings and a missed release when the tool was built. Each episode now runs through six commands: ingest the master, transcribe locally with word timestamps, assemble (trim the dead air, add the intro and outro, normalize to broadcast loudness, export video and audio with a shifted caption file), break down with a language model into chapters, clip picks, Short picks, a teaser, titles, a description, tags, and pre-publish flags, cut the pieces snapped to sentence boundaries with a vertical version of each Short, and pack the metadata for every platform. Every piece gets a package slot and a publish time derived from the release date; the queue page opens on that package as a strip; the calendar exports the posts and the production milestones with two alarms each. The Shorts step runs through Higgsfield's Model Context Protocol server from a Claude session.

What to look at

  • The package strip: twenty days, one cell per post, the state of each piece at a glance, today outlined.
  • Trimming: nudge an in or out point by a quarter second and re-cut, snapped to a sentence boundary from the transcript.
  • The deliverables table: a spare swapped into a taken slot moves the other piece out, and the calendar follows.

Numbers

8
pieces in the package, on a fixed fourteen-day rhythm counted
90 s
for a 36-minute episode through the local speech model on an M1 Max measured
~3 h
of human time per episode after the first run; the rest runs unattended stated in the handover, not yet measured
2
alarms per event in the exported calendar counted

Boundaries

Built for a real client; the show on this page is fictional, with synthetic episodes and placeholder clips, and none of the client's guests or channel figures appear. The language-model and Shorts steps need keys; transcription, assembly, cutting, and packing run locally without any.

See it yourself

bin/bcb gui opens the desk; bin/bcb doctor checks the tools. The redesign lives on the redesign/production-desk branch.

Grad Cam Control

A software-defined multi-camera production rig, documented so someone else can run the show.

In one minute

A software-defined, network-controlled multi-camera production system for long live events: mixed-generation cameras on network lens-control nodes, a phone camera on a robotic gimbal sending video over the network, a switcher, show playback, and streaming, driven from one Mac. Version two is a full revision of version one; the analysis lists every issue found and how it was fixed, including killing the riskiest item โ€” reverse-engineering a gimbal โ€” in favour of a supported interface. Safety is designed in: every continuous motion carries a roughly 250-millisecond lease so a dead network can never leave a lens moving; state is read back, not assumed; official interfaces over reverse engineering. The event-day runbook is written to be executable by someone who did not build the system.

What to look at

  • The runbook's T-minus checkpoints and the readiness go or no-go script.
  • The decision record that replaced a reverse-engineered protocol with a supported one.
  • The day-of note that tables exactly what was real and what was not.

Numbers

29
commits counted
13
test files, simulator-driven counted
14
numbered design documents plus decision records counted
250 ms
motion lease, the fail-safe counted

Boundaries

As of 2026-08-29 the system had not carried a live ceremony; the only proven camera path was the phone node, and the note says so.

See it yourself

The simulator and tests run locally; the runbook, the decision records, and the day-of note read in two minutes.