I had a working fractal renderer in Go. Three fractal types, interactive zoom and pan, targeting 800×600 at 60 fps. The rendering loop spawned 600 goroutines per frame, one per scanline, each iterating every pixel through the escape-time algorithm and writing RGBA bytes into a shared buffer. It worked fine on the desktop. Then I compiled it to WebAssembly and those 600 goroutines became cooperative tasks on a single browser thread1. The parallelism evaporated. Single-digit frame rates.

Each pixel is independent, with no data dependencies between neighbouring pixels. This is exactly the kind of work a GPU is built for. Ebiten has a shader language called Kage, so I moved the escape-time calculation there. The resulting pipeline is strange enough to be worth documenting: Go source compiles to WASM for the host logic, while Kage shaders compile at runtime to WebGL for the per-pixel work.

The CPU Version Link to heading

The original render loop:

var wg sync.WaitGroup
wg.Add(h)
for y := 0; y < h; y++ {
    go func(row int) {
        defer wg.Done()
        for x := 0; x < w; x++ {
            cx, cy := vp.ScreenToComplex(x, row)
            iter := ft.Iterate(cx, cy, maxIter)
            c := colormap.Map(iter, maxIter)
            offset := (row*w + x) * 4
            pix[offset] = c.R
            pix[offset+1] = c.G
            pix[offset+2] = c.B
            pix[offset+3] = c.A
        }
    }(y)
}
wg.Wait()

At 800×600, that is 480,000 independent calculations per frame. The cost of each one depends on how long its coordinate takes to escape, up to the configured iteration limit.

I measured the renderer on my Framework 13 with an i7-1360P and Intel Iris Xe graphics. The native figures are the mean of three Go benchmark runs. The browser figures come from Chromium 149 over five seconds of alternating left and right input, which invalidates the CPU cache on every frame without walking away from the fractal.

Native CPU benchmark256 iterations1024 iterations
Time per redraw~11.5 ms~36.1 ms
WASM browser frame rate256 iterations1024 iterations
CPU~19.8 fps~8.0 fps
GPU, scalar path60 fps60 fps
GPU, double-single at 1e-1060 fps~42.8 fps

The browser is capped at 60 fps by display synchronisation, so those GPU rows describe visible frame cadence rather than shader execution time. Ebiten’s reported draw-call time only measures CPU submission because the GPU work completes asynchronously. The paired arithmetic is substantially more expensive at 1024 iterations, but it still leaves the deep-zoom renderer usable. Press G in the embed to compare it with the CPU path.

Kage Link to heading

Ebiten’s shader language is called Kage2. It looks like Go but compiles to GLSL, HLSL, or MSL depending on the platform. On WASM, it becomes WebGL shader code.

Kage has restrictions that bite you immediately:

  • No dynamic loop bounds. The for loop needs a hardcoded upper limit compiled into the shader.
  • There is one floating-point type, float, rather than separate float32 and float64 types. The Kage specification does not guarantee its precision; on the WebGL implementations I tested, it behaves as float32.
  • No recursion.
  • Host values such as viewport state enter through global uniform variables.

My shader uses Kage’s three-argument fragment entry point: Fragment(dstPos vec4, srcPos vec2, color vec4) vec4. You get a pixel position and return a colour. The GPU schedules one invocation for each of the 480,000 pixels.

The escape-time loop in Kage:

for i := 0; i < 10000; i++ {
    if float(i) >= MaxIter {
        break
    }
    zx2 := zx * zx
    zy2 := zy * zy
    if zx2+zy2 > 4.0 {
        escaped = true
        break
    }
    zy = 2.0*zx*zy + cy
    zx = zx2 - zy2 + cx
    iterCount += 1.0
}

The 10000 is a compiled-in ceiling. MaxIter, passed as a uniform, is the actual limit. The GPU compiles this into a fixed-iteration loop with early exit. Wasteful in theory; irrelevant in practice because the GPU is doing 480,000 of these simultaneously.

For colouring, I use the continuous iteration count: iter + 1 - log2(log2(|z|²)) / 2. This maps escape iterations to a smooth float value rather than discrete integers, eliminating the harsh colour banding that makes cheap fractal renderers look cheap. The colour gradient uses Bernstein polynomial basis functions, same formula as the CPU version so you can toggle between renderers mid-session without a visible jump.

Precision Link to heading

The CPU renderer uses float64 throughout. The shared viewport clamps MinZoom at 1e-14, where the visible coordinate range spans fourteen decimal places. Float64 has 15-16 significant digits, so the CPU path retains enough precision.

On the WebGL implementations I tested, Kage’s float behaves as float32: about seven significant digits. Pass the viewport centre as a single float uniform, compute each pixel’s coordinate by adding a sub-pixel offset and once you zoom past about 1e-6, adjacent pixels begin mapping to the same value. The fractal freezes into a static grid. You’re zooming into a point the shader can no longer subdivide.

Double-Single Emulation Link to heading

I split each float64 viewport coordinate into two float32 values on the Go side.

func Split(value float64) (high, low float32) {
    high = float32(value)
    low = float32(value - float64(high))
    return high, low
}

high is the float32 approximation. low captures the remainder lost in conversion. Within the viewport’s coordinate range, reconstructing the pair in float64 stays within 1e-15 of the original value.

Keeping that precision in the shader requires every operation to return another pair. I represent a double-single number as a vec2, with the high component in x and the low component in y. Addition uses an error-free transform and renormalises the result instead of collapsing it back to one float:

func dsNormalize(hi, lo float) vec2 {
    s := hi + lo
    e := lo - (s - hi)
    return vec2(s, e)
}

func dsAdd(a, b vec2) vec2 {
    s := a.x + b.x
    v := s - a.x
    e := (a.x - (s - v)) + (b.x - v) + a.y + b.y
    return dsNormalize(s, e)
}

Multiplication needs the same treatment. The constant 4097 is 2¹² + 1, which splits a float32 significand into two parts so the rounding error from a.x * b.x can be recovered.3

func dsMul(a, b vec2) vec2 {
    p := a.x * b.x

    aCon := a.x * 4097.0
    aHi := aCon - (aCon - a.x)
    aLo := a.x - aHi
    bCon := b.x * 4097.0
    bHi := bCon - (bCon - b.x)
    bLo := b.x - bHi

    e := ((aHi*bHi - p) + aHi*bLo + aLo*bHi) + aLo*bLo
    e += a.x*b.y + a.y*b.x + a.y*b.y
    return dsNormalize(p, e)
}

func dsSub(a, b vec2) vec2 {
    return dsAdd(a, vec2(-b.x, -b.y))
}

func dsScale(a vec2, b float) vec2 {
    return dsMul(a, vec2(b, 0))
}

The deep-zoom path keeps cx, cy, zx and zy as pairs through the entire escape loop. Converting back to a scalar before iteration would throw away the extra precision I had just recovered.

cx := dsAdd(vec2(CenterHigh.x, CenterLow.x), vec2(offsetX, 0))
cy := dsAdd(vec2(CenterHigh.y, CenterLow.y), vec2(offsetY, 0))
zx := vec2(0, 0)
zy := vec2(0, 0)

for i := 0; i < 10000; i++ {
    zx2 := dsMul(zx, zx)
    zy2 := dsMul(zy, zy)
    if zx2.x+zx2.y+zy2.x+zy2.y > 4.0 {
        break
    }
    zy = dsAdd(dsScale(dsMul(zx, zy), 2.0), cy)
    zx = dsAdd(dsSub(zx2, zy2), cx)
}

The shader only activates this more expensive path below 3.5e-7. At 800 pixels wide, that is roughly where the per-pixel coordinate delta starts approaching float32’s least significant bit. Above this zoom level, scalar arithmetic is faster and precise enough. Below it, the high/low pairs keep adjacent pixels distinct.

The practical limit lands around 1e-10 before artefacts creep back in. Not as deep as the CPU’s 1e-14, but enough for this viewer.

Wiring the Renderer Link to heading

Ebiten can’t compile shaders until an active graphics context exists, which means the first frame. Try to compile during init() or NewGame() and you get a nil pointer deep in the OpenGL backend. The GPU renderer lazy-initialises on the first Update() call, falls back to the CPU path if compilation fails, and from there passes viewport state as uniforms each draw call:

func (sr *ShaderRenderer) Draw(screen *ebiten.Image, opts *RenderOpts) {
    cxHi, cxLo := precision.Split(opts.CenterX)
    cyHi, cyLo := precision.Split(opts.CenterY)

    uniforms := map[string]interface{}{
        "CenterHigh": []float32{cxHi, cyHi},
        "CenterLow":  []float32{cxLo, cyLo},
        "Zoom":       float32(opts.Zoom),
        "MaxIter":    float32(opts.MaxIter),
        "Resolution": []float32{float32(opts.Width), float32(opts.Height)},
    }

    drawOpts := &ebiten.DrawRectShaderOptions{Uniforms: uniforms}
    screen.DrawRectShader(opts.Width, opts.Height, sr.shaders[sr.currentIndex], drawOpts)
}

No pixel buffer. No goroutines. No sync.WaitGroup. One DrawRectShader call and the GPU handles the rest.

Shipping It Link to heading

wasm:
    GOOS=js GOARCH=wasm go build -ldflags="-s -w" -o web/fractal.wasm ./cmd/fractal/
    cp -f "$$(go env GOROOT)/lib/wasm/wasm_exec.js" web/

The resulting binary is 12 MB. That’s Go’s WASM story: the entire runtime, garbage collector and goroutine scheduler ship in the binary regardless of whether you use them. The -s -w flags strip symbols and DWARF, which saves about 3 MB compared to a debug build. The deploy gzips it to just under 3 MB over the wire. Lazy loading is not optional at that size.

The wasm_exec.js glue file must come from the same Go toolchain release that compiled the binary. A version mismatch produces silent failures or cryptic Go.run errors with no useful stack trace. Ask me how I know.

Ebiten’s WASM target manages its own canvas element, which makes embedding it inside a blog post slightly awkward. It insists on owning the full document. I therefore run the viewer in a minimal iframe rather than letting it take over the article. This is document containment, not a security boundary: the viewer is code I built and serve from the same origin, so the site trusts it in the same way it trusts its other JavaScript. Ebiten’s own example gallery also uses iframes to contain its demos.

The embed is lazy-loaded behind a click. A 12 MB binary is too much to fetch speculatively on page scroll.

Try It Link to heading

Click inside the viewer to give it keyboard focus. Scroll to zoom, drag to pan and press G to switch between the CPU and GPU renderers.

Loading WebAssembly…

  1. Go’s WASM target runs goroutines cooperatively on a single OS thread. There’s no true parallelism; goroutines yield at scheduling points but never run simultaneously. See the Go WebAssembly wiki for details on the execution model. ↩︎

  2. Kage documentation. Kage (影) means “shadow” in Japanese. ↩︎

  3. The paired arithmetic uses Dekker splitting and Knuth’s two-sum algorithm. See Dekker, T.J. (1971) “A floating-point technique for extending the available precision” and Priest, D.M. (1991) “Algorithms for arbitrary precision floating point arithmetic.” ↩︎