Hit testing without a DOM
The interface on this site is a picture painted on a curved mesh. Working out what you clicked takes six coordinate transforms and a formula that has to exist twice.
The interface on this site is built from three.js meshes and text rendered into a RenderTexture, mapped onto a curved CRT model, and pushed through a shader that bends it. There are no elements in it anywhere. When you click something in there, no part of the browser knows what you clicked.
event.target is a DOM concept and there is no DOM. There's a canvas, and inside it a mesh, and painted on that mesh a picture of an interface.
Getting from a click to "they pressed the close button on the third window" is a coordinate problem with about six steps in it, and every step is one you normally never have to see.
The chain
Here's the actual projection, lightly trimmed:
const project = (clientX: number, clientY: number) => {
const box = canvas.getBoundingClientRect()
probe.ndc.set(
((clientX - box.left) / box.width) * 2 - 1,
-((clientY - box.top) / box.height) * 2 + 1,
)
probe.ray.setFromCamera(probe.ndc, mainCamera)
const [hit] = probe.ray.intersectObject(mesh)
if (!hit) return null
const vu = (SCREEN.zMax - hit.point.z) / (SCREEN.zMax - SCREEN.zMin)
const vv = (hit.point.y - SCREEN.yMin) / (SCREEN.yMax - SCREEN.yMin)
const cx = vu * 2 - 1
const cy = vv * 2 - 1
const bend = 1 + CURVATURE * (cx * cx + cy * cy) * 0.16
const ux = cx * bend * 0.5 + 0.5
const uy = cy * bend * 0.5 + 0.5
return {
u: (ux - SAFE_INSET.x) / (1 - 2 * SAFE_INSET.x),
v: (uy - SAFE_INSET.y) / (1 - 2 * SAFE_INSET.y),
}
}Client coordinates become normalized device coordinates. A ray goes out from the camera. If it misses the glass there's no hit and the event is rejected, which is the whole answer to "did they click the screen at all". If it lands, the world-space point becomes a UV by measuring it against the screen's bounds, and then the curvature and the safe-area inset get applied.
The result feeds two things. A compute function hands it to the portal's raycaster as normalized device coordinates, so R3F's normal event system keeps working inside the texture. A pick function converts it to OS pixels for anything that needs a coordinate rather than an event.
Two implementations of one curve
The shader bends the image. The picker has to bend the pointer by exactly the same amount, or the cursor lands somewhere other than where the thing it's pointing at appears to be.
So bend exists twice. Once in GLSL, deciding which texel to sample. Once in TypeScript, deciding which element you hit. Two implementations of one formula in two languages, and nothing checks that they agree.
Picking through a distortion
Move the slider away from 1.60 and the picker stops agreeing with what you can see.
Drag that slider. Near the center almost nothing happens, because the distortion is smallest there. The corners go wrong first, which is the worst available failure mode: it works while you're testing in the middle of the screen and breaks for the controls you put at the edges.
What event.target was doing
Computing hits by hand makes it obvious how much the DOM was handling.
- Testing the topmost element first, then working down
- Respecting z-order, including stacking contexts
- Not hitting anything an ancestor has clipped out of view
- Pointer capture, so a drag keeps receiving events after the cursor leaves the element
- Focus, tab order, and keyboard activation
- Bubbling, so a handler on a container catches clicks on its children
R3F returns some of it. Raycast intersections come back sorted by distance, so z-order is free as long as your layers really do sit at different z positions. The rest you write.
Pointer capture goes stale
Dragging a window was the first thing that broke. Pointer capture inside the portal doesn't refresh event.point, so a drag moves a few pixels and then freezes while the cursor keeps going.
The fix is to stop using synthetic events for drags. On pointer down, attach native listeners to the window and run every move through pick() for a fresh coordinate:
const onPointerDown = () => {
const move = (event: PointerEvent) => {
const point = pick(event.clientX, event.clientY)
if (point) setPosition(point)
}
window.addEventListener('pointermove', move)
window.addEventListener('pointerup', () => window.removeEventListener('pointermove', move), {
once: true,
})
}This is also what makes a drag keep working once the cursor leaves the canvas, which the synthetic path won't do at all.
Clipping happens in two coordinate systems
A window body has to clip its contents. Meshes clip with three.js clipping planes, which are world space. Text is troika, which clips with a clipRect in the text object's own local space.
Two systems, two frames, one window. The local-space one is where the mistakes live.
A clip rect in the wrong frame
The clip rect tracks the scroll offset, so the rows stay inside the window.
Scroll that with the offset unaccounted for and the rows walk straight out of the window. The clip rect was correct at offset zero, which is exactly the state you'd check it in.
The version that works keeps the scroll offset in context and has every label subtract it before setting its rect. Getting it wrong doesn't throw and doesn't warn. It renders text where text shouldn't be, or clips it away completely, which reads as the text failing to render.
The one that took longest
The OS lays out at the screen's aspect ratio, except in fullscreen, where it lays out at the viewport's.
The clipping component imported the screen aspect constant directly instead of reading the live value. In the normal view those are the same number, so it worked. In fullscreen every window clipped about 120 pixels too far to the left.
It didn't present as a clipping bug. It presented as selection bars and row icons sometimes not rendering, because those were the things living in the leftmost 120 pixels. I spent a while reading the components that drew them.
Anything converting OS pixels into world or UV space has to read the live metrics. A constant that's usually right is worse than one that's always wrong, because the always-wrong one gets caught the first time you run it.
Worth the trouble
The OS renders inside the texture so the CRT shader applies to it: barrel, scanlines, tube mask, phosphor bleed. A DOM overlay approximating that in CSS was tried and rejected, because an approximation of the distortion isn't the distortion.
What that costs is a hand-written hit test, a formula maintained in two languages, two clipping systems, and drag handling that skips the framework's event system entirely.
event.target covers all of it in one property, and I had never once thought about it before I had to build it.