How to Judge a Hackathon Fairly: A Complete Guide 2026

To judge a hackathon fairly, score every team against the same published rubric with several independent judges, recuse anyone with a conflict of interest, give each team equal time, and combine scores with a tie-break rule you announce before the event. Fairness comes from the process, not from the panel’s good intentions. Most unfair outcomes trace back to a rubric that was never written down, time limits that quietly drifted, or a decision rule nobody published.

The practical version takes about two weeks of organizer work and 60 to 90 minutes on judging day. I’ve sat on both sides of this, and the events that feel fair afterward all share the same traits: the criteria went out with the call for entries, the scorecard was identical for every team, and the results included the arithmetic.

Here is the process I would use, in the order I would use it.

Table of Contents

What You Need

What You Need

Build the framework before applications arrive, not the night before judging. Everything below is a decision you will otherwise have to make in a noisy room with tired people.

  • Competition goals. What is this hackathon actually for? If a city is looking for deployable pilots, your criteria must say so. If it is looking for technical range, they must say that instead. Mixed goals produce mixed scores.
  • Challenge criteria. The 5 to 7 things you will actually score, drawn from the challenge brief.
  • Scoring weights. What share each criterion carries, totalling 100.
  • Submission requirements. Repo link, two-minute video, deployed URL, slide deck. State what happens when a team cannot demo live.
  • Judge roles. Who scores, who observes, who is the tie-break chair, and who runs the timer.
  • Conflict-of-interest rules. Who counts as conflicted and what recusal looks like in practice.
  • Time limits. Presentation length, demo length, question length, and how many teams a judge sees.
  • Tie-breaking procedure. Written down before any score exists.
  • A feedback template. Two or three sentences every team receives, winners included.

If any of those nine is missing, the fairness gap is already open. Judges will fill the blank with whatever they personally value that day, and different judges fill it differently.

Step-by-Step: How to Judge a Hackathon Fairly

1. Set the judging criteria before judging begins

Turn the challenge brief into a short list of observable criteria. Observable means two judges watching the same pitch would write down similar evidence. “Innovative” is not observable on its own; “uses a method no other team attempted this weekend” is.

A workable set for a civic or smart-city hackathon: user need, technical feasibility, responsible use of data, real-world impact, originality, user experience, and pitch clarity. Seven is already the practical ceiling. Beyond that, judges start trading precision for speed.

Write the rubric into the call for entries. Teams build toward a published brief far more effectively than toward a criteria list revealed on demo day, and judges score more consistently when everyone arrived knowing the target.

2. Write a scoring rubric with measurable anchors

Each criterion needs a weight and a description of what a 1, a 3 and a 5 look like. Without anchors, judges quietly score on confidence, design taste and how senior the team looks. With anchors they score evidence.

Example weighted rubric for a smart-city mobility hackathon
CriterionWeightA 1 looks likeA 3 looks likeA 5 looks like
User need25%Solves a problem nobody statedReal need, named usersReal need with evidence from actual users
Technical feasibility20%Static mockup onlyCore path runs end to endRuns with error handling and a deploy path
Responsible data use20%Personal data handled looselySources named, consent impliedDocumented consent, minimisation, no personal data stored
Real-world impact15%Vague improvement claimOne deployment path identifiedPartner named and willing to pilot
Originality10%Template or tutorial buildNew approach within a known patternApproach nobody else on the day attempted
User experience10%Judge cannot complete the flowCore flow works, rough edgesClear flow someone unfamiliar can follow

Two judges using this grid on the same submission should land within one point of each other on any criterion. Where they do not, the fix is usually a vague anchor, not a bad judge. Note that responsible data use can carry a fifth of the weight in civic events and almost nothing in a gaming jam. That is not inconsistency, it is the criteria matching the challenge.

3. Standardize submissions and presentation time

Equal time is the cheapest fairness win available. Set a presentation length, a demo length, and a question window, then run them with a visible timer and a two-minute warning.

Decide in advance how you handle three situations. A demo that crashes gets a recorded backup video, not an automatic deduction. A team that finishes early does not get bonus attention. A team that arrives with fifteen slides and no working build gets judged on the same criteria as everyone else, which usually means losing technical-feasibility points rather than getting a sympathy adjustment.

Accessibility accommodations get granted quietly and in advance, never announced to the panel. Extra time to present because a team member needs it should not change what the judges are measuring.

4. Assign judges and check conflicts of interest

Build a panel that covers the challenge rather than one made of the loudest available names. A smart-city event usually needs a civic or policy voice, someone who can assess technical execution, someone who speaks for the affected users, and a community or nonprofit perspective.

Send every judge one short form before the event: name, organization, any team they know personally, any team they mentored or employed, and any sponsor they work for. Recusal is a normal part of judging, not an accusation, and judges should know that in writing.

Plan the replacement. If a recusal lands mid-event, the substitute comes from a standby list rather than the room, and the substitute scores the withdrawn team’s slot on the same rubric. Recusal is the normal case rather than the exception, so budget for it like any other part of the schedule.

5. Score independently so panel pressure cannot reach the scorecard

Every judge completes and submits their scorecard privately before any discussion happens. No consensus in the room, no show of hands, no pointing at the loudest project.

This one rule removes several biases at once. Halo effects from a confident pitch fade when the score is written before anyone agrees. Anchoring on the first strong or weak team stops because each judge meets the submissions in a different order. Popularity effects cannot accumulate in the open.

Rotate presentation order or randomize it per judge so order effects do not land on the same teams every time. Judges report anchoring and off-rubric drift most often on the last teams of the day, which is exactly what order randomization prevents.

6. Run a structured calibration discussion

Before real scoring, give every judge the same one or two sample submissions and have them score them alone. Then compare distributions and talk about differences.

Ask about evidence, assumptions, and what each judge needed but did not see. When two judges are three points apart, someone has usually misread an anchor, and that is worth naming in front of the panel. First-time judges tend to score wildly until a practice round pulls them together, and no prior judging experience is required to serve, which is exactly why calibration matters.

7. Deliberate using an agreed decision rule

Combine independent scores first and discuss second. The commonest aggregation rule is the weighted mean across judges, which is easy to explain and hard to argue with. Drop the highest and lowest score per criterion if you have five or more judges per team, which limits the damage of one outlier.

Decide your tie-breakers before scores exist. A workable order: highest weighted mean, then highest score on the criterion carrying the most weight, then a recorded discussion vote among judges with no conflict, then a chair’s casting vote.

Not every disagreement is real. Separate arithmetic and interpretation differences from substantive ones, so ten minutes of deliberation are not spent re-adding columns. Write down the rule you used and the reason for any override. If a team asks how they lost, you can answer with a method instead of an apology.

8. Give every team useful, respectful feedback

Convert the scorecard into three or four sentences: the criterion that carried the most weight, one specific strength with the evidence for it, one risk they should address next, and the single change that would most raise their score.

Feedback goes to every team, before or alongside the winner announcement, and it is tailored. A generic script reads as indifference. Deliver losing feedback separately from the results if you can, so nobody hears their project critiqued as a prize announcement.

Judging an online or async hackathon

Remote judging adds two problems: video quality and unequal working conditions. A quiet room with a good microphone is an advantage that has nothing to do with the project.

Review submissions blind where you can, with team names and links hidden until scores are submitted. Require video from the same fixed prompt, keep the same length for every team, and set a single time window rather than rolling deadlines. Score the repository and the running app directly, not just the recording, and judge recorded material the same way you would judge live material.

FactorIn-person demo dayOnline or async
EvidenceLive demo plus pitchRunning app plus recorded pitch
Order effectRandomize per judgeRandomize per judge
Q and ATwo minutes, liveWritten questions, same window
Blind reviewOptionalStrongly recommended
Judges per team3 to 53 to 5

Time budgeting decides whether any of this holds at scale. With 100 teams at 4 minutes each you have over six hours of presentation, and no judge can do that well. Stagger judges across parallel rooms so each judge sees 25 to 30 teams, cap expo-style judging at about 90 seconds per project, and promote a documented shortlist rather than letting one exhausted panel decide everything.

Community consensus is consistent on the number: at least three independent judges per team for any result you intend to defend.

Common Mistakes

  • Vague criteria. Fix: write anchors for 1, 3 and 5 and test them on a sample submission before the event.
  • Rules that change mid-event. Fix: freeze the rubric 48 hours before judging and send any clarification to all judges and all teams at once.
  • Unequal presentation time. Fix: visible timer, two-minute warning, and no bonus attention for teams that finish early.
  • Judging on polish. Fix: cap pitch weight at 10 or 15 and put technical execution and real-world impact above design quality.
  • Undisclosed conflicts. Fix: conflict form before the event, recusal without penalty, standby judge on call.
  • Group discussion before scoring. Fix: private scorecards submitted first, discussion opened only after.
  • Inconsistent scoring between judges. Fix: calibration round on shared samples, and compare distributions before any deliberation.
  • Post-hoc criteria. Fix: add new criteria to the rubric for the next event, never to this one.
  • Failing to explain the decision. Fix: publish the aggregation rule, tie-breakers and a short reason for each award.
  • Template builds and faked demos. Fix: ask teams to state what they built this weekend, check commit history, and score templates under originality with no penalty for a modest but honest build.

On AI tooling, write the rule down before the event: whether assistants are permitted, what they may be used for, and whether teams must declare it. Then apply it identically, including to the winning team.

Frequently Asked Questions

How many judges does a hackathon need?

Three to five independent judges per team is the working minimum for results people will accept. Fewer than three means one strong personality decides the outcome, and one tired judge at the end of the day can swing several projects. For 100 or more teams, stagger judges across parallel rooms so each judge sees roughly 25 to 30 projects rather than scoring everything badly.

Should hackathon judges know the team names?

Ideally not, at least until scorecards are submitted. Hiding names, affiliations and pitch order removes halo effects from confident speakers and well-funded teams, and it costs organizers little since submissions can be listed by number. For in-person events, blind review is harder, so randomized order and private scorecards become the substitute.

How do you break a tie in hackathon judging?

Decide the tie-breaker chain before any scores exist and publish it with the criteria. A common order: highest weighted mean, then highest score on the criterion carrying the greatest weight, then a recorded vote among judges without a conflict, then a casting vote from a named chair. Agreeing on the chain in advance is what keeps a tie from becoming an argument.

How much weight should business value carry?

Match the weight to the challenge, and keep it under 20 unless the event explicitly exists to find commercial products. Civic, social-impact and smart-city events usually reward feasibility, responsible data use and a named deployment partner instead. Weighting business value heavily in an impact event rewards pitch polish over the work the brief actually asked for.

What do you do when a submission raises a safety or privacy concern?

Flag it outside the scoring flow. Ask the organizer to pull a named reviewer with the relevant expertise onto that submission, and record the concern as a condition of any award rather than as a deduction applied by one judge alone. In a civic hackathon this usually means personal data handling, and it deserves a written check before a winner is announced.

Can you use ChatGPT in a hackathon?

Rules differ by event, so state yours before the hackathon starts. Most events permit assistants for research, scaffolding and documentation while requiring teams to declare substantive use. The judging question is narrower than the policy question: when every team has the same tooling, judge what they built and verified, not whether they used an assistant at all.

Conclusion

Start with the one thing you can do this week: write the criteria, the weights, the 1-3-5 anchors and the tie-break chain into a single page and publish it with the call for entries. Everything after that is smaller.

Collect conflict forms a week ahead, run a 30-minute calibration round on shared samples, and require private scorecards before any conversation about winners. Then answer every team with their scores, the method you used, and something specific they can act on. That process can be tested after the event and improved next time, which is the whole point: fair judging is a design decision, not a personality trait. Updated for 2026.

Leave a Comment