Gitframes

Hacker News by 22 min read 63x views
Gitframes

Share Post

gitframes — code-first video, rendered natively on WebGPU

Photoshop-, After Effects-, and Blender-class video tools as one npm bundle that AI agents run alongside code.

npm license status discord youtube node engine gpu vision

⚠️ Beta: gitframes is under energetic development. APIs may alter between releases and several features may be incomplete or unstable.

Gitframes is built for coding agents. It packs the activity group normally divided throughout three desktop apps (Photoshop-grade compositing and VFX, After Effects-style motion, typography and keyframing, and Blender-style 3D scenes, cameras and models) into one lightweight npm package. Your delegate writes a TypeScript composition, checks frames, and renders an MP4, and nobody has to instal or licence a multi-gigabyte imaginative suite.

Code-first video as clean application engineering — no headless browser, no DOM reflow, no screenshot pipeline. Renders immediately on GPU hardware via Dawn / WebGPU / Metal / Vulkan in Node.js and contemporary WebGPU browsers.

Every example of these films is rendered by gitframes from TypeScript in examples/. Click a motionless to observe it on YouTube.

Note

Using an AI coding agent? Install the gitframes skills in one line.

Claude Code

/plugin instal gitframes 

Any another agent (Codex, Cursor, Hermes, Gemini CLI, Copilot, and more)

npx skills add gatewai-dev/gitframes

See Agent Skills & Plugins for details.



Modern automated video generation is normally constrained by the architectures of general-purpose web browsers: procedure overhead, non-deterministic DOM layout reflows, and dilatory screenshot capture. Gitframes treats video construction as application engineering:

Pillar What it means
🚀 Zero Headless-Browser Overhead No Puppeteer, no Chromium IPC, no page.screenshot(). Gitframes talks direct to native GPU devices via Dawn/WebGPU and hardware-encodes alongside @napi-rs/webcodecs.
🎯 Deterministic Frame-Accurate Clock Absolute example clocks, discrete example points, and frame-accurate audio BeatGrids. No floating timers, no drift, no dropped frames.
🔠 Analytic, Resolution-Independent Type The Slug algorithm evaluates glyph contours per-pixel in WGSL — no texture atlases, no scaling artifacts, razor-sharp from 10 px to 10,000 px.
🎨 Photoshop-Grade Tonal & Spatial VFX 50+ standardized GPU shaders: Curves, Levels, Selective Color, 3D LUTs, Halftone, Film Grain, Unsharp Mask, Mesh Warp, and Screen-Space Relighting.
🧊 Unified 3D & 2D Depth Compositing Nest 2D flex/box trees inner 3D homography planes, multiplane rigs, and meshes (OBJ, FBX, glTF/GLB, STL, PLY, VOX, 3DS, OFF), alongside PBR pane and SSAO.
🔊 Built-in Procedural Audio DSP Multi-track soundtracks, deterministic procedural passage SFX (whoosh, impact, riser), and reactive signals that run visuals from audio.
👁️ On-Device Neural Vision Object tracking, case segmentation, multi-person pose, and individual mattes from Apache-2.0 ONNX models — feeding reactive signals without a circular trip to disk.
☁️ Cloud-Native & CI/CD Ready ~200–400 MB RAM per employee (vs. 2–4 GB for Chromium), ideal for serverless GPU render clusters (AWS G4/G5, Modal, RunPod, Kubernetes).

Architectural Comparison: Gitframes vs. Remotion vs. Hyperframes

Developers generating video programmatically commonly measure Remotion (React/Chromium) or Hyperframes (Canvas2D/SVG web animation). The matrix below compares the essential engineering dimensions.

Detailed Comparison Matrix

Capability / Dimension Gitframes Remotion Hyperframes
Underlying Engine Native WebGPU (WGSL compute & render pipelines via Dawn / Metal / Vulkan) Chromium / Puppeteer (React DOM, HTML/CSS layout) Canvas2D / WebGL / SVG (browser or Node Skia)
Rendering Architecture Direct hardware framebuffer rendering & hardware video encoding (@napi-rs/webcodecs) Spawns headless Chrome; captures frames via CDP / page.screenshot() Software or hardware 2D canvas context
Throughput 60–120+ FPS (real-time to faster-than-real-time GPU execution) 5–20 FPS (DOM reflow, IPC, rasterization) 20–40 FPS (CPU diagram commands / JS)
Memory Footprint ~200–400 MB per render (zero browser) 1.5–4.0 GB+ per employee (Chromium + V8 DOM heap) ~500 MB–1 GB (Skia/Canvas bindings)
Typography Engine Slug GPU — analytic Bézier evaluation in WGSL, infinite zoom, After Effects selectors Browser DOM content (CSS fonts, rasterized, blurry under 3D transforms) Canvas2D / way content (CPU-rasterized glyphs)
2D VFX & Post-Processing 50+ WebGPU shaders (Curves, Levels, Selective Color, 3D LUT, Film Grain, Halftone, Liquify, PBR Glass, Relight) CSS Filters or tradition WebGL canvas wrappers Basic Canvas2D composites and 2D filters
3D Graphics & Depth Native 3D environment graph — LookAt/Turntable camera, multiplane, skinning (OBJ/FBX/glTF), SSAO, PCSS, DoF None built-in (embed Three.js/Fiber inner React DOM) Minimal 2.5D layers; no unified mesh pipeline
Motion Blur & Physics Physical 180° shutter speed buffers in MRT + closed-form fountain kinematics CSS transitions / JS interpolation; synthetic blur hacks Frame interpolation or manual multipass
Audio Engine & DSP Native audio DSP & procedural SFX (multi-track mixing, attack grids, reactive signals) <Audio> playback; essential quantity curves Basic fixed audio playback
Charts & Data Viz Layer.chart — line, area, bar, scatter, candlestick, pie and donut charts built from native vector nodes, alongside staggered disclose animations DOM diagram libraries (Recharts, Chart.js) Custom canvas diagram operations
AI & Computer Vision On-device ONNX vision — COCO-80 finding + case masks (RTMDet-Ins), COCO-17 stance (RTMO), individual mattes (Selfie Segmenter); WebGPU tensor conditioning (Canny, depth-to-normals, optical flow, deflicker) External pre-rendered assets; no native GPU tensor conditioning External pre-rendered assets
Headless Verification FrameGrid communication sheets, single-frame snapshots, Skia MSE pixel-invariant assertions Playwright/Puppeteer ocular snapshots Manual example inspection / canvas diffing
Docker / Cloud Portability Compact (~500 MB slim depiction alongside native GPU/Vulkan drivers) Heavy (~2–3 GB alongside Chromium, fonts, X11/Mesa) Moderate receptacle size

Key Features & Capabilities

1. Slug GPU Vector Typography & AE Parity

Traditional content relies on CPU rasterization or low-res SDF atlases that soften under 3D camera sweeps. Gitframes integrates the Slug algorithm (SlugPipeline):

  • Analytic GPU evaluation — WGSL part shaders resolve exact cubic/quadratic Béziers per-pixel. Glyphs remain keen at 10 px or 10,000 px alongside zero CPU re-rasterization.
  • After Effects–parity animators — range selectors (square, ramp_up, ramp_down, triangle, smooth), easeHigh/easeLow curves, and seeded PRNG character shuffling (TextAnimator).
  • Human typing cadence — valued symbols delays (commas 3×, declaration ends 5.5×, newlines 7×) and trailing scramble resolution (TypewriterAnimator).
  • 3D volumetric formations — map content onto cylindrical drums, logarithmic vortex spirals, and double-helix ribbons alongside surface-normal banking (evaluateVolumetricFormation).
  • Dynamic foremost & skew — area-preserving unimodular shear and accordion line-leading anchored to baseline, center, or top.

2. Photoshop-Grade WebGPU 2D VFX (50+ Shaders)

A thorough suite of expert image/video shader nodes in nodes/ and packages/webgpu-renderers:

  • Tonal grading — Curves (RGB/R/G/B spline), Levels (black/white point, gamma, output), Shadows/Highlights, Selective Color (CMYK gamut isolation), 3D Cube LUT (ApplyLUT).
  • Stylization & grain — Film Grain (Gaussian emulsion alongside spatial kernel variation), Halftone (mono/RGB/CMYK, adjustable dot form & angle), Gradient Map, High Pass.
  • Optics & lens — Bilateral Gaussian Blur, Unsharp Mask, Vignette, Refraction Caustics, PBR Glassmorphism alongside colorful dispersion (PBRGlass).
  • Distortion & warping — Displacement Maps, Liquify, Mesh Warp, Corner Pin homography.

3. Unified 3D Scene Graph, Camera & Mesh Shading

  • Calibrated camera rig — LookAt and Turntable cameras (Camera3D) calibrated so z = 0 matches 2D canvas pixel coordinates 1:1.
  • 3D layout primitives — Layer3D.cube, carousel, prism, plane, grid alongside unified depth-buffer testing.
  • Zero-dependency example parsers — OBJ, FBX, glTF/GLB, STL, PLY, VOX, 3DS, OFF.
  • Skeletal animation & shading — 128-bone Linear Blend Skinning, Blinn-Phong & PBR multi-light shading, PCSS/Poisson communication shadows, SSAO, and optical DoF.
  • Physical motion blur — 180° shutter motion blur alongside per-vertex speed vectors packed into rg16float MRT buffers.

4. Audio Layers, Procedural SFX & Reactive Signals

  • Soundtrack layers — .audio media nodes alongside frame-exact lifecycle control.
  • Procedural SFX — deterministic CPU-synthesized whooshes, impacts, risers, downshifters, and glitches placed on the bar/beat grid (renderSfx, mixSfxInto, softLimit).
  • Multi-track mixing — expert tracks headlessly alongside mixAudioTracks and encodeStereoWav.
  • Reactive signals — run transforms, scale, borders, or shader uniforms from tempo signals (Signal.builder) or audio analysis.

Layer.chart builds line, area, bar (grouped or stacked), scatter, candlestick, pie and donut charts. d3 computes the scales, ticks and geometry; all bar, line, piece and tag is an average box, way or content node:

  • Labels use the composition's registered fonts and the identical GPU content renderer as the remainder of the film.
  • A built-in disclose draws lines on, grows bars from the baseline and staggers points and slices (animate: { start, duration, stagger, comfort }, or animate: false).
  • The diagram is one box, so it positions, animates, grades and tilts into 3D akin any another layer.
Layer.chart( { type: "bar", width: 900, height: 480, categories: ["Q1", "Q2", "Q3", "Q4"], series: [ { name: "Revenue", data: [12, 19, 24, 31] }, { name: "Costs", data: [8, 11, 13, 15] }, ], yAxis: { format: "$,.0f" }, animate: { start: 10, duration: 30 }, }, { position: "absolute", x: 120, y: 200 }, );

6. On-Device Vision & Tracking

@gitframes/vision runs ONNX models via onnxruntime-node (CPU) or onnxruntime-web (WebGPU) and wires all outcome into the identical reactive indication exterior the remainder of Gitframes consumes.

Tip

Lazy by construction. VisionRunner.create(), comp.withVision(...) and VisionNode.attach(...) execute zero I/O — no downloads, no sessions, no document probes. A example is fetched the archetypal period a project really runs. To heated up onward of time, call await runner.preload(["detect", "pose"]) (or await vision.ready() on an attached node).

Every example is Apache-2.0, pinned to an immutable Hugging Face revision, and verified by SHA-256 following download.

Task Option Model Output
Detect enableDetection RTMDet-Ins t/s/m (OpenMMLab) COCO-80 boxes + scores, tracked complete time
Segment enableSegmentation RTMDet-Ins (same onward continue as detect) Soft per-instance masks, frame-aligned
Pose enablePose RTMO t/s/m (OpenMMLab) 17 COCO keypoints + visibility per person
Matte enableMatte MediaPipe Selfie Segmenter (Google) Fast person-vs-background alpha for picture / webcam framing
  • Variants — variant: "t" | "s" | "m" (default "s"; ~24 / 43 / 116 MB for RTMDet-Ins). Tune confidence and a COCO classes display per composition. On CPU, a 2K example takes approximately 200–340 ms to detect + segment, ~120 ms for stance and ~20 ms for the matte alongside "s".
  • One pass, two tasks — finding and segmentation portion a sole RTMDet-Ins conclusion per frame.
  • Picking a matte — the Selfie Segmenter is tuned for a individual filling much of the frame: it misses distant figures and can study "person" on close-ups alongside nobody in them. For item else, cut out alongside case masks (matteSource: "instance", the default).
  • Whole-subject cutouts — disguise / matte / crop modes merge all comparably sized case that overlaps the chief subject, so a flowing attire or a held device stays attached to the person, during a tunnel or opening framing them does not.
  • One-frame delay — imagination says all layer's former rendered frame, so results trail the dish by one example and example 0 has none. Verify imagination layers alongside the exported video or successive frames, not example grids.
  • Model cache & mirrors — models are cached atomically (temp + rename) in $GITFRAMES_MODELS_DIR (default ~/.cache/gitframes/models). Point baseUrl or GITFRAMES_MODELS_BASE_URL at your own mirror for air-gapped or CI renders.

Temporal tracking & analysis

  • Multi-object tracker (TemporalObjectTracker) assigns stable trackIds via IoU association, alongside configurable minHits, positionSmoothing, and velocity-based coasting for up to maxMissedFrames (default 15) so a transient young female holds the track alternatively of flashing.
  • Pose↔track matching (pose-track-matcher) binds keypoints to the correct track by id, afterward by spatial IoU fallback.
  • One-shot sequence analysis — comp.analyzeVisionSequence(src, { tasks, categories }) decodes frames through the mediabunny pipeline, tracks them, and returns a zod-serializable study (per-track example ranges, average speed, sampled center paths, per-class presence/confidence, average disguise coverage, example download bytes/timing) (analyzeSequence).

Every tracked being is exposed as reactive ProgrammaticSignals that enliven layers and shader uniforms:

Group Highlights
objects get(trackId), byCategory(cat, rank), primary, count, hasCategory, detectedCategories
objects.*.bounds x/y/width/height, screenX/screenY/screenWidth/screenHeight, aspectRatio, area
objects.*.anchors 9 anchors (corners, edges, center) prepared for pinning
objects.*.kinematics vx, vy, speed, acceleration, headingRad/Deg
objects.*.pose All 17 COCO keypoints, affirmative hasPose, wristSpeed, handRaised, bodyTiltAngle
masks get(trackId), subject, count; per-mask area, coverage, solidity, bboxFill
segmentation subject, humanSilhouette, instanceMasks, matte.coverage, GPU stencilTexture
classes Per-class count, maxConfidence, present, primary, affirmative a finding histogram
Tensors poseLandmarksTensor [17,3], objectsTensor [16,8], masksTensor [16,2], histogramTensor [80]

Project normalized landmarks to display area alongside a configurable camera FOV, afterward tie any node to a track or landmark (SpatialLandmarkTransformer, spatial-pin):

  • pinToObject(track, { anchor, offsetX/Y/Z, matchWidth, matchHeight, smoothFrames, hideWhenLost })
  • pinToLandmark(coord, { offsetX/Y/Z })

High-level construction helpers

  • Subject Sandwich — comp.addSubjectSandwich({ source, behind, feather, fit }) cuts the foreground topic out and places typography/graphics rearward them.
  • Smart Reframing — comp.addSmartFraming({ source, target, targetAspect, damping, leadHeadroom }) auto-crops 16:9 → 9:16 during tracking target.
  • Subject Outline — comp.addSubjectOutline(vision.segmentation.subject, { source, color, width, blur }) strokes the segmented border as an audio-reactive contour glow.
  • Tracked Region Blur — layer.blurRegion(track, { power }) blurs faces, plates, or any detected class.
  • Node modes — passthrough, mask, matte, crop, skeleton, boxes, tracking; choice the cutout alpha alongside matteSource: "instance" | "selfie", and optionally keyBackground to develop the topic into connected foreground.
  • Runtime config is zod-validated and accessible from a zod-only entry (@gitframes/vision/schemas) so the hot way stays zod-free. Unknown or removed options are rejected, not silently ignored.
  • vision.summary(frame) returns a deterministic, serializable snapshot (objects, classes, masks) harmless to call inner a example hook.
  • Clear failures — a example that is the incorrect size, fails its checksum, or lacks an expected output raises an error naming the example and its source.
  • Browser entry — @gitframes/vision/web re-exports the motor affirmative createWebGPUProvider() / hasWebGPU(); onnxruntime-web is an optional lazy peer.

7. Headless Conformance & FrameGrid Testing

  • Pixel-sampling invariant assertions — test compositions in Vitest alongside skia-canvas to verify shader math, font coverage, and Mean Squared Error (MSE) temporal deltas.
  • FrameGrid communication sheets — comp.renderFrameGrid(...) outputs sequential-frame communication sheets for instant assessment of easing, kinetic type, and transitions.

8. Live Preview in the Browser

  • Runs your composition, not a video — startPreview({ entry, export }) serves a localhost WebGPU participant that loads the composition's own component and renders all example live in the browser. Nothing is streamed: the server lone hands complete the bundle, the project's assets, and the soundtrack blended by the export engine.
  • Timeline, waveform & example stepping — play/pause, scrub, stage example by frame, and peruse resolution, FPS, duration, and audio position at a glance.
  • One stable URL per project — the harbor is derived from the operating directory, so re-running the preview replaces the operating server and any open tab reloads into the new type by itself. Close the tab and the server shuts downward concerning five seconds later.
  • Shown anywhere you are — startPreview serves the leaf and returns its URL alternatively of beginning a browser, so an delegate can display it in its own pane (Claude Code, Codex); continue open: true to open the scheme browser.
import { startPreview } from "gitframes"; const session = await startPreview( { entry: new URL("./film.ts", import.meta.url), export: "buildFilm" }, { title: "gitframes film" }, ); console.log(`Preview at ${session.url}`); await session.closed; // serves until its tab closes or a newer preview takes over

41133


Managed alongside pnpm workspaces and turbo:

gitframes/ ├── packages/ │ ├── gitframes/ # Unified SDK (Composition, Layer, LayerAnimation, Signal, effects) │ ├── core/ # Core AST, Effect basis class, VirtualMediaData, imagination types │ ├── compositions/ # Layout engine, Flex/Box AST compiler, timeline evaluator │ ├── webgpu-renderers/ # WGSL shaders, Slug content engine, 3D renderer, camera, lights, materials │ ├── tensor-webgpu/ # WebGPU compute pipelines (Canny, depth-to-normals, flow, deflicker, landmarks) │ ├── vision/ # ONNX imagination engine: detect, segment, pose, matte, tracking, signals │ ├── renderer/ # Headless Node.js WebGPU renderer via Dawn, WebCodecs, skia-canvas │ ├── renderers/ # Higher-level render orchestration │ ├── media/ # Media decoding / encoding adapters │ ├── node-sdk/ # Node renderer contracts and outcome schemas │ ├── server-utils/ # Server infrastructure, storage, asset caches │ └── client-utils/ # Shared browser utilities ├── nodes/ # 58+ specialized domain nodes (VFX, audio, layout, node-vision) ├── apps/ │ └── renderer-service/ # Production HTTP / gRPC rendering microservice container ├── examples/ # Reference compositions and films ├── plugins/gitframes/ # Agent plugin: skills lone (setup, compose, effects, render) └── scripts/ # Build, release, and plugin validation tooling 

Requirements: Node.js ≥ 22. Gitframes uses native GPU acceleration via Dawn / WebGPU or Vulkan.


1. Basic Composition & Kinetic Auto-Layout

import { Composition, Layer, LayerAnimation } from "gitframes"; // 1. Initialize a 1080p60 composition const comp = new Composition({ width: 1920, height: 1080, fps: 60, durationFrames: 180, // 3 seconds backgroundColor: "#090a0f", fonts: ["assets/fonts/Inter.ttf", "assets/fonts/SpaceGrotesk.ttf"], }); // 2. Define bodily snap-overshoot animations const cardEntrance = LayerAnimation.create() .fadeIn(0, 20, "power2.out") .fromTo("y", 60, 0, { start: 0, end: 35, ease: "back.out(1.5)" }) .fromTo("scale", 0.92, 1.0, { start: 0, end: 35, ease: "back.out(1.2)" }); // 3. Assemble a receptive flex-layout card const heroCard = Layer.box({ width: 720, height: 380, background: "#141721", borderRadius: 24, borderColor: "#262b3d", borderWidth: 1.5, padding: 32, children: [ Layer.flex({ dir: "column", gap: 16, children: [ Layer.text("GITFRAMES ENGINE", { fontSize: 16, fontWeight: 700, fill: "#6366f1", letterSpacing: 2.0, }), Layer.text("Next-Gen WebGPU Motion", { fontSize: 48, fontWeight: 700, fill: "#f8fafc", fontFamily: "SpaceGrotesk", }), Layer.text("Direct hardware video construction without headless browser overhead.", { fontSize: 20, fill: "#94a3b8", lineHeight: 28, }), ], }), ], }).animate(cardEntrance); comp.add(heroCard);

2. Unified 3D Scene alongside Camera & 3D Model

import { Composition, Layer, Layer3D, CameraAnimation, Light } from "gitframes"; const comp = new Composition({ width: 1920, height: 1080, fps: 60, durationFrames: 300 }); // 1. LookAt 3D camera alongside a uninterrupted orbit const cameraAnim = CameraAnimation.camera().orbit({ azimuth: { from: -30, to: 30 }, elevation: { from: 15, to: 15 }, radius: { to: 1200 }, start: 0, end: 300, }); comp.add( Layer.camera({ x: 960, y: 540, z: -1000, targetX: 960, targetY: 540, targetZ: 0 }).animate(cameraAnim) ); // 2. Studio lighting comp.add(Light.ambient("#ffffff", 0.4)); comp.add(Light.directional({ color: "#e0e7ff", intensity: 1.2, x: 500, y: -800, z: -600 })); // 3. 3D example alongside skeletal animation comp.add( Layer.glb("assets/models/character.glb", { x: 960, y: 640, z: 0, scale: 2.5, material: "lit", loop: true, }) ); // 4. 3D prism layout carousel comp.add( Layer3D.carousel({ radius: 400, items: [ Layer.box({ width: 280, height: 180, background: "#1e293b", borderRadius: 16 }), Layer.box({ width: 280, height: 180, background: "#334155", borderRadius: 16 }), Layer.box({ width: 280, height: 180, background: "#0f172a", borderRadius: 16 }), ], }) );

3. Audio Soundtrack, Procedural SFX & Reactive Signals

import { Composition, Layer, LayerAnimation, Signal, renderSfx, mixSfxInto, softLimit } from "gitframes"; const comp = new Composition({ width: 1920, height: 1080, fps: 60 }); const totalFrames = 240; // 1. Soundtrack layer comp.addAudio(Layer.audio("assets/score.mp3", { volume: 0.9, durationFrames: totalFrames })); // 2. Frame-accurate procedural SFX on the attack grid const bed: [Float32Array, Float32Array] = [ new Float32Array(Math.ceil((totalFrames / 60) * 48000)), new Float32Array(Math.ceil((totalFrames / 60) * 48000)), ]; mixSfxInto(bed, [ renderSfx({ type: "whoosh", atBar: 0.79, volume: 0.5 }, { sampleRate: 48000, secondsPerBar: 2.0, seed: 1 }), renderSfx({ type: "impact", atBar: 1.0, volume: 0.8 }, { sampleRate: 48000, secondsPerBar: 2.0, seed: 2 }), ]); softLimit(bed); // 3. Tempo indication (120 BPM = 2 Hz) const beatPulse = Signal.builder({ type: "sawtooth", frequency: 2, amplitude: 0.08, offset: 1.0 }); // 4. Bind it to visuals const reactiveCard = Layer.box({ width: 400, height: 250, background: "#1c202e", borderRadius: 20 }) .animate( LayerAnimation.create() .signal("scale", beatPulse, { multiplier: 1.0, offset: 0.0 }) .fromTo("opacity", 0, 1, { start: 0, end: 15, ease: "power2.out" }), ); comp.add(reactiveCard);

4. Chained WebGPU Post-Processing VFX

import { Composition, FilmGrain, Vignette, ColorBalance } from "gitframes"; const comp = new Composition({ width: 1920, height: 1080, fps: 60 }); // Whole-composition cinematic class + movie emulsion comp.apply(new Vignette({ strength: 0.28, radius: 0.85 })); comp.apply(new FilmGrain({ strength: 0.06, size: 1.5, animated: true })); comp.apply( new ColorBalance({ shadows: { cyanRed: 0, magentaGreen: 2, yellowBlue: 6 }, highlights: { cyanRed: 4, magentaGreen: 1, yellowBlue: -2 }, }), );

5. Vision: Pin, Matte & Reframe

import { Composition, Layer, Vignette } from "gitframes"; const comp = new Composition({ width: 1920, height: 1080, fps: 30 }); // Run imagination on the entire composition. Models download lazily on archetypal use. const vision = comp.withVision({ enableDetection: true, enableSegmentation: true, enablePose: true, variant: "s", confidence: 0.35, }); // Pin a caption to the chief tracked topic (smoothing + auto-hide whenever lost) comp.add( Layer.text("SUBJECT 01", { fontSize: 40, fill: "#f8fafc" }).pinToObject( vision.objects.primary, { anchor: "topCenter", offsetY: -48, smoothFrames: 5, hideWhenLost: true }, ), ); // Drive a shader uniform from a reactive indication — here, topic disguise coverage comp.add( Layer.box({ width: 1920, height: 1080, background: "#000000" }).withEffect( new Vignette({ strength: vision.segmentation.subject.coverage, radius: 0.9 }), ), ); // Or use the one-liners for the average editorial moves: // comp.addSubjectSandwich({ source: "assets/dancer.mp4", behind: [headline], feather: 4 }); // comp.addSmartFraming({ source: "assets/action.mp4", target: vision.objects.primary, targetAspect: 9 / 16 }); // comp.addSubjectOutline(vision.segmentation.subject, { source: "assets/character.mp4", color: "#FF5A1F", width: 6 }); // Inspect a origin before authoring: one-shot, ffmpeg-free, zod-serializable report const report = await comp.analyzeVisionSequence("assets/street.mp4", { tasks: ["detect", "pose"], categories: ["person"], }); console.log(report.tracks.map((t) => `${t.category}#${t.trackId} ${t.frames.join("–")}`));

Standalone runner (no composition):

import { VisionRunner } from "@gitframes/vision"; const runner = VisionRunner.create({ variant: "s", confidence: 0.3 }); // zero I/O const frame = { data: rgba, width: 1920, height: 1080 }; const boxes = await runner.detect(frame); // downloads RTMDet-Ins on archetypal call const { masks } = await runner.segment(frame); // identical onward pass, no second inference const { group } = await runner.pose(frame); // RTMO, COCO-17 keypoints runner.close();

In the browser (WebGPU EP):

import { VisionRunner, createWebGPUProvider, hasWebGPU } from "@gitframes/vision/web"; if (hasWebGPU()) { const runner = VisionRunner.create({ provider: createWebGPUProvider() }); }

6. Headless Video & FrameGrid Rendering

import { buildMyComposition } from "./my-composition.js"; const comp = await buildMyComposition(); // 1. Single example to a PNG buffer for ocular inspection const frameBuffer = await comp.renderFrame({ frame: 45 }); // 2. Contact-sheet grid of 12 sequential frames const gridBuffer = await comp.renderFrameGrid({ startFrame: 0, endFrame: 120, stepFrames: 10, cellWidth: 320, showLabels: true, }); // 3. Final hardware-encoded MP4 alongside blended audio const { filePath } = await comp.renderVideo({ outputPath: "output/final-product-film.mp4", quality: "high", concurrency: 4, }); console.log(`Video rendered successfully to: ${filePath}`);

Engineering Doctrines & Best Practices

  1. Design tokens & theme contracts — define a centralized THEME for colors, type, radii, and spacing. Never hardcode magic hex values or ad-hoc margins.
  2. WebGPU premultiplied-alpha invariant — part shaders outputting premultiplied alpha (color * opacity * alpha) must use srcFactor: "one" in their blend province ({ srcFactor: "one", dstFactor: "one-minus-src-alpha", operation: "add" }). Never use srcFactor: "src-alpha" for premultiplied output — squaring alpha darkens fades into murky gray.
  3. Carrier equivalent cuts — transport a ocular component (badge, card, cursor, container) throughout environment boundaries alongside uninterrupted speed and stance to evade jarring cuts.
  4. Physical easing vocabulary — back.out(1.4–1.7) for snap-overshoot entrances, fountain / expo.out for decelerating motion, power2.in for exits. Reserve linear for infinite spinners and period counters.
  5. Headless invariant verification — verify shader transforms, glyph coverage, and temporal MSE deltas alongside skia-canvas pixel sampling in Vitest before shipping.

Gitframes ships delegate skills that instruct Claude, Codex, and another coding agents how to write, render, and inspect compositions. The plugin (gitframes) is listed in Anthropic's authoritative plugin directory and contains only skills — no MCP servers, hooks, or commands. Every another delegate gets the identical skills through the skills CLI.

Skill Use it for
gitframes Starting a project: instal from npm, scaffold a construction and render script, archetypal verified render
gitframes-compose Compositions, tier trees, layout, animation and easing, attack grids, movie structure
gitframes-effects Effect classes, the unified division architecture, premultiplied-alpha invariants, imagination conditioning
gitframes-render Headless rendering, FrameGrid inspection, pixel probes, MP4 shipment checks

Once installed, skills burden automatically whenever a project matches (e.g. "add a film-grain continue to this scene" or "render a example grid of intro.ts").

What the plugin runs and sends

The plugin is instructions only. It bundles no executables, MCP servers, hooks, or bundle launchers, and it sends no data anywhere. The skills inform your delegate to add the gitframes npm bundle to your project and how to use it. When that code uses on-device vision, the SDK downloads the pinned example weights from Hugging Face on archetypal use (see On-Device Vision). Nothing alternatively leaves your machine.

/plugin instal gitframes 

Or from your shell:

claude plugin instal gitframes@claude-plugins-official

It installs from Anthropic's authoritative marketplace, which Claude Code adds for you, so there is no market step, and plugins from it update automatically. Afterwards, restart Claude Code or run /reload-plugins. /plugin commands need an interactive claude terminal; in the desktop app's Code tab, use the casing form or + > Plugins > Add plugin and choice Gitframes.

Add --scope project to the casing form to document the plugin in .claude/settings.json for the entire team.

Enable it for everyone in your repo. Commit this to .claude/settings.json; Claude Code prompts teammates to instal it whenever they rely the folder:

{ "enabledPlugins": { "gitframes@claude-plugins-official": true } }

Straight from this repository (tracks chief alternatively of the directory release):

/plugin market add gatewai-dev/gitframes /plugin instal gitframes@gitframes-plugins 

Codex, Cursor, Hermes, and another agents

The skills CLI installs the skills into any of 70+ agents, including Codex, Cursor, Hermes, Gemini CLI, GitHub Copilot, Windsurf, OpenCode, and Goose:

npx skills add gatewai-dev/gitframes

It detects the agents on your device and asks anywhere to install. To choose them yourself, continue -a formerly per agent, add -g to instal for your person alternatively of this project, and -y to skip the prompts:

npx skills add gatewai-dev/gitframes -a codex -a cursor -a hermes-agent -g -y

Keep them current alongside npx skills update, and eliminate them alongside npx skills remove.

Or copy the folders by hand: put plugins/gitframes/skills/<name>/ into .claude/skills/, .agents/skills/, or ~/.agents/skills/. VS Code / Copilot / Cursor / Kiro can burden the portable plugin.json through their plugin UI.

The plugin lives in plugins/gitframes/ so installs transport lone the skills; users get the motor from npm. Two manifests there depict it: plugin.json (portable Agent Plugins 1.0, which additionally carries the OpenAI listing metadata) and .claude-plugin/plugin.json. The market catalog is .claude-plugin/marketplace.json. The portable site set is closed — client-specific sectors go in that client's manifest, not in plugin.json. The type in the two follows the gitframes package: pnpm run version:packages syncs it following changeset type (or run pnpm run sync:plugin-version on its own), since clients use it to decide whenever to update.

Inside this repository, Claude Code and another agents choice up skills through the symlinks in .agents/skills/ and .claude/skills/. Skills live lone under plugins/gitframes/skills/; never copy them elsewhere. pnpm run check:plugins validates manifests, accomplishment frontmatter, market catalogs, symlinks, and the generated effects catalog. pnpm run sync:effects-catalog regenerates the gitframes-effects catalog following any Effect category change.


Reference Showcase Examples

The examples/ directory holds production-grade citation compositions:


Gitframes uses pnpm (10+) and turbo for orchestration.

# Install pnpm install # Build all packages pnpm build # Run conformance tests pnpm test # Check the imagination models end to end (downloads ~380 MB of weights once) pnpm --filter @gitframes/vision test:models # Render a particular showcase example pnpm --filter @gitframes/example-21-full-circle render # Render the expert brand film cd examples/19_gitframes_film && pnpm render

Docker Container for Production Rendering

An optimized Dockerfile.renderer deploys the renderer assistance into haze GPU clusters:

docker build -t gitframes-renderer -f Dockerfile.renderer .


Gitframes is open-source application licensed under Apache-2.0. The imagination models it downloads on petition — RTMDet-Ins and RTMO (OpenMMLab) and the Selfie Segmenter (Google) — are additionally Apache-2.0; see registry.ts for exact sources and checksums.

Other Article Hacker News
↑
Close Right Ads
Close Left Ads