TK. Talal Khawaja, home Résumé

Case study · 2026 · Hackathon

CaptionForge AI

Turns a short video clip into four caption styles by sampling frames in the browser and calling Fireworks models, with honest fallbacks and human review before anything is posted.

AMD Developer Hackathon: ACT II (Track 2) · lablab.ai · 13 Jul 2026

CAFIG. 01 — SYSTEM SKETCHGENERATED FROM STACK · NOT A SCREENSHOTNEXT.JSTYPESCRIPTFIREWORKS AIDOCKER
Fig. — generated system sketch from the project’s stack. Not a product screenshot.

Overview

CaptionForge AI turns short video clips into four caption styles: formal, sarcastic, humorous-tech and humorous-non-tech. It was a six-person teqprotech entry for Track 2 of the AMD Developer Hackathon: ACT II on lablab.ai. The team was Umer Anis, Talal Khawaja, Aqeela Urooj, Shadab Akhund, Muhammad Yousuf Maqbool and Darren Melvern.

Problem

A clip for social media often needs more than one caption. You might want a straight one, a wry one, one for a technical audience and one for everyone else. Writing four voices by hand is slow. Handing it to a model has its own risk: the model can speak confidently about audio it never heard or details it never saw, and a demo can quietly keep going on canned output when the model is down. We wanted captions that were quick to produce and honest about where they came from.

How it works

The user uploads a short video and can add context or a transcript. In the browser, an HTML video element and a canvas pull up to five downsized JPEG preview frames, which are kept in React state. When the user clicks Generate Captions, the frontend sends the frames, the optional context, the filename and the duration to a Next.js API route. The route validates the request, limits the frame count and payload size, and calls Fireworks through the OpenAI-compatible SDK.

There are three tiers:

  1. Fireworks Vision. The configured vision model reads the frames and writes all four styles.
  2. Fireworks Text. If the vision model returns an unavailable status, the route immediately tries a text model. That fallback uses only the user's context or transcript, the filename and the duration. It doesn't pretend to analyse audio, voice, identity or frames it never received, and the UI shows a warning that vision was unavailable.
  3. Mock. If both Fireworks paths fail, the app returns safe mock captions so the demo can continue, and labels them as mock.

A separate one-shot Docker agent handles the track's judging format. It reads tasks from a JSON file, samples each video with FFmpeg, calls the deployed caption API, writes a results file and exits. The image holds no secrets and never starts the web server. The interactive app itself runs on Vercel.

Key decisions and trade-offs

  • Label the source every time. Each caption card carries a badge saying Fireworks Vision, Fireworks Text or Mock, so nobody mistakes a fallback for real analysis.
  • Degrade honestly. The text fallback only claims what it was given.
  • Keep limits out of the captions. Technical caveats go in a visual summary and safety note, not inside the caption text a user might copy.
  • Humans post. Captions are meant to be reviewed before publishing. The README warns that AI can miss context, speech, sarcasm or sensitive content, and that the app avoids naming private people unless the user supplies the names.

Problem

Writing captions for short videos in several voices is repetitive, and AI captioning can quietly overreach. It may claim to understand audio or details it never saw, or keep going when a model is down without saying so.

Approach

The browser samples up to five small frames from the uploaded clip. A server-side route sends them, with optional context, to a Fireworks vision model through the OpenAI-compatible SDK. If the vision model is unavailable, it falls back to a text model that uses only the supplied context, filename and duration. If both fail, it returns labelled mock captions. Every result shows which path produced it.

My contribution

Team member (teqprotech).

  • Umer AnisTeammate
  • Talal KhawajaTeam member (teqprotech)
  • Aqeela UroojTeammate
  • Shadab AkhundTeammate
  • Muhammad Yousuf MaqboolTeammate
  • Darren MelvernTeammate

Architecture

01Next.js02TypeScript03Fireworks AI04Docker
Diagram — the recorded stack, in project-record order. A sketch, not a screenshot or a data flow.

The detailed architecture for CaptionForge AI hasn’t been documented yet, so this sketch only lists the technologies on the project record. Nothing here is guessed.

Features

  • Browser-side frame sampling with HTML video and canvas

  • Four styles: formal, sarcastic, humorous-tech, humorous-non-tech

  • Fireworks Vision first, then Fireworks Text, then labelled mock captions

  • Source badge on every caption card

  • Visual summary and safety note alongside the captions

  • API keys kept server-side

  • One-shot Docker judging agent using FFmpeg

Stack

Shipped on lablab.ai, GitHub and Vercel.

Outcome

Submitted to AMD Developer Hackathon: ACT II (Track 2) on lablab.ai, 13 Jul 2026.

Result not recorded.

Built with Teqprotech.