Skip to main content
SYS_INIT // DWARKESH.LAB
DWARKESHRAMANI
Computer Engineering · Full-Stack Engineer
INITIALIZING DWARKESH // BUILD SYSTEM...12%
02/AI Engineer & Backend Architect·2026·SHIPPED

AI Hackathon Judge

Multi-persona autonomous consensus judge and multimodal project evaluator.

Python 3.11FastAPIGoogle Gemini 2.5OpenCVAsyncIODockeryt-dlpPydantic

01 // OVERVIEW

Multi-persona autonomous consensus judge and multimodal project evaluator.

THE PROBLEM

Hackathon judging is notorious for human fatigue, unconscious bias, inconsistent rubrics, and shallow inspections where flashy slides overshadow non-functional code or leaked production credentials.

THE IDEA & APPROACH

An autonomous multi-persona evaluation system executing parallel reviews (The VC, The CTO, Product Manager, UI/UX Designer, CS Professor) to aggregate scores mathematically while deeply inspecting code trees, live DOM weights, and video presentations.

02 // SYSTEM ARCHITECTURE & DATA FLOW

Parallel asynchronous ingestion engine executing static repository AST traversal, live DOM asset scans, 4-tier video transcript fallback, and multi-agent synthesis.

FIG 1.0 // MULTI-PERSONA PARALLEL CONSENSUS MATRIXSUBMISSIONGitHub + DOM + VideoRegex Secret SweepFASTAPIasyncio.gatherParallel OrchestratorTHE VC (Market/TAM)THE CTO (Code Quality)PRODUCT MGR (UX/Fit)UI/UX DESIGNERCS PROFESSOR (Rigor)SCORE MATRIXWin Probability™ RubricHonest Reality Check
DETAILED EXECUTION SEQUENCE
01 →User submits repository link, live demo URL, pitch deck (PDF/PPTX), and demo video
02 →BFS file crawler parses repository dependency manifests and computes entry point LOC
03 →Security regex engine sweeps tree for exposed AWS keys, GitHub tokens, and private keys
04 →4-layer video pipeline extracts audio transcripts via client-side fetch, cookies, or Gemini Vision
05 →asyncio.gather executes 5 persona prompts in parallel against unified project context
06 →Mathematical aggregator computes Win Probability™ and generates brutal 'Why You Won't Win' critique

03 // TECHNICAL DECISIONS & TRADE-OFFS

DECISION5 distinct domain personas rather than a single prompt
ALTERNATIVES CONSIDEREDMulti-Judge Modeling
WHY

A single prompt produces homogenized scores. Distinct personas unmask conflicting trade-offs (e.g., CTO scores high for code, VC scores low for TAM).

DECISION4-tier fallback (Browser client -> Server cookies -> yt-dlp -> Gemini Vision)
ALTERNATIVES CONSIDEREDVideo Analysis Resiliency
WHY

Cloud hosting platforms (Render, AWS) suffer frequent IP blocking from YouTube. Multi-layer fallback guarantees 99.8% ingestion success.

04 // CHALLENGES & RESOLUTIONS

CHALLENGE // 01

Mitigating context window exhaustion when ingesting 50k+ LOC repositories by implementing smart token pruning.

CHALLENGE // 02

Eliminating single-judge hallucinations by computing cross-persona variance scores.

CHALLENGE // 03

Extracting structured rubrics from messy PPTX and PDF slide decks without losing speaker notes.

05 // VERIFIED OUTCOMES

Deployed and running live on Render with comprehensive rubric breakdown across 6 scoring vectors.

Identifies credential vulnerabilities and fake prototype claims with 94%+ precision.

Provides actionable, candid pre-pitch feedback for student and hackathon teams.

06 // LESSONS & TAKEAWAYS

Multimodal analysis must fail gracefully: if a demo video is blocked, the code and pitch deck analysis must continue smoothly.
Persona-driven critique produces drastically higher qualitative value for builders than generic praise.
STATUS: VERIFIED ON GITHUB