Introducing INKBOT
An experiment in “intelligence architecture,” human intent, and making AI output more inspectable before it becomes trusted.
Thanks for the context and examples, @forum-helper. I wanted to share an independent prototype I have been developing called INKBOT.
It began with a specific problem I repeatedly encountered as a builder: the friction between what a person naturally means and what a multimodal AI system infers. Models are excellent at producing raw outputs, but getting to a clear, accurate, human-meaningful interpretation often becomes a separate alignment problem.
Instead of asking users to learn prompt engineering, INKBOT acts as an intermediate intelligence architecture layer that structures the human meaning between human intent and AI output. Its working loop is:
https://ko-fi.com/thomascoates/shop
I have put together an evaluation package containing complete runnable local HTML files for both core versions of the prototype:
INKBOT Lite 71
The entry-level abstraction. It takes about five minutes to test: users describe concepts in ordinary language, receive a usable visual handoff, and can inspect whether the first picture reflects what they actually meant.
The Intelligence Architecture Edition
This sits beneath the simple interface. It explores what happens when a user wraps a model’s generation loop in explicit programmatic structures such as concept memory, references, provenance validation, local database retrieval, revision histories, and version histories.
What I am testing
One of the practical case studies included in the package involves a complex shore-shell mapping workflow: a field technician, a survey wheel, sequential photography, GPS mapping coordinates, and overlapping frames. INKBOT successfully translated that multi-step workflow into a clean, unified system diagram that can be inspected visually before execution.
The visual below is not intended as a claim that the implementation is complete or production-ready. It is a concrete, inspectable artifact of the kind of human-to-AI translation problem I am exploring.
Suggested Additional Visuals for the Evaluation Package
I am expanding the package with additional visual evidence so reviewers can inspect the workflow from multiple angles rather than relying on product claims alone.
I am asking the community to critically inspect the design rather than simply accept the framing. I would welcome feedback from people working on multimodal systems, instruction following, model behavior, evaluation, human–AI interaction, safety, or agentic workflows.
In particular: where does this duplicate existing work, where does the architecture become genuinely useful, what important limitations am I missing, and where would experienced builders simplify or challenge the design?
Questions I would value feedback on
Thanks for taking the time to look at an independent community project.