LLM's generate one token at a time, but we usually consume them as turns - that's missing most of the fun! Logit Loom turns the token boundary into a programmable surface. Play with the stream, mess with logits, observe exactly what gets admitted, and get right inside the generation loop.
I’ve open-sourced a Rust toolkit for building precise, inspectable generation pipelines around llama.cpp: compose ordered logit transforms, observe exact token bytes, stop at defined boundaries, checkpoint sessions, and retain serializable receipts of what ran.
It’s a foundation for custom sampling and steering, constrained generation, agent runtimes, reproducible inference experiments, and the mechanics layer of neuro-symbolic systems. It's not going to write the fun bits for you, but it will take the tedium OUT of writing those bits.
Be your own Dr. Frankenstein and play with the nascent minds of your agentic friends!