Fake Work vs Real Progress: The Rubric We Grade Founders Against
The actual rubric GRILLR uses to grade founder work: five tests, three evidence tiers, three red flags, and the hard caps that make unproven work impossible to score well.
fake work
Most founders do not fail because they stopped working. They fail because they spent four weeks doing things that felt productive and proved nothing.
Redesigning the logo. Rewriting the deck. Reading one more book about customer discovery. All of that is work. None of it is evidence.
GRILLR grades every task a founder submits on a 1-to-10 scale, and the grader is built around one question: is there proof, or is there only a claim? This article publishes the actual rubric behind that score. You can apply it to your own work this week without signing up for anything.
The rule the entire rubric is built on
A bare link proves nothing. A file name proves nothing. "I reached out to some people" proves nothing.
Our grader is explicitly told that it cannot open links, so a URL on its own carries no weight. It can only judge what the written description and the attached files make concretely verifiable. That constraint is deliberate. It forces the standard a skeptical investor already applies to you: show me the thing, do not tell me the thing exists.
Before any score is assigned, the grader must answer five questions in a fixed order, and it is forbidden from jumping ahead to the number. That sequencing matters more than the score, because it forces the work to be examined before it is judged. Most self-assessment fails precisely here: founders decide how they feel about the week first, then hunt for reasons.
Test 1: What did this task actually require?
The first step is not looking at your work at all. It is writing down, concretely, what "done" demanded, and what proof would show it was truly done.
Real proof looks like a working link, a screenshot of real output, actual numbers, named people you contacted, a published page, or a reply from a real human being.
Doing this first is what stops the most common self-deception in early-stage building: quietly redefining the task to match whatever you happened to produce. You set out to talk to ten potential customers, you talked to two friends, and by Friday the task has become "gathered initial feedback." Write the requirement down before you look at the output, and that move becomes impossible.
Test 2: What evidence is actually there?
Now the grader quotes the real, concrete proof present in the submission: the actual links, the specific numbers, the real names, exactly what a screenshot shows. If there is nothing but claims, this field is filled in with a single word — none.
Evidence then gets sorted into exactly three tiers:
- Strong — the proof clearly shows the work was done.
- Weak — the proof is vague, thin, or unverifiable.
- None — there is no real proof, only assertions.
There is one important exception, and it exists because rigid rubrics punish honest work. Research, verification, and decision tasks often produce no natural artifact. Checking a subreddit's rules, comparing three suppliers, picking a target segment — none of these generate a screenshot. For those, a specific and checkable written account counts as real evidence: exact names, the exact rule or fact you found, and where you found it. "I checked, it seems fine" does not count. "r/solopreneur bans self-promotion under rule 3, so that channel is out" does.
Test 3: What is the single biggest gap?
Next the grader names the one most important thing missing or wrong versus what the task required. Not a list. One thing.
Forcing a single answer prevents the most useless kind of feedback, the scattergun list of twelve minor improvements that leaves you with no idea what to do on Monday morning. There is always one gap that matters most. Naming it is the whole job.
This test also contains the rubric's most important piece of fairness. If you deviated from the prescribed steps but genuinely engaged with them and can give a concrete reason for the substitution, that is an adaptation, not off-topic work. You get graded on the work you actually did plus the quality of the substitute, and the gap becomes whatever is still unverified. Founders who adapt intelligently to real constraints should not score the same as founders who ignored the task.
Test 4: The three red flags
Three specific patterns are flagged before scoring, and any one of them caps your grade at 3 out of 10:
- The submission is entirely a promise. Pure future tense with nothing done yet. "I will contact them this week" and nothing else.
- Restating the task instead of doing it. Describing the work in more detail is not the work.
- Work unrelated to what was asked. Effort spent elsewhere is still effort spent elsewhere.

Two clarifications keep this from being unfair. A future-tense next step written alongside completed work is not a red flag, it is planning. And a reasoned adaptation is not off-topic work, as covered above.
If you want a single diagnostic for your own week, use the first one. Read back what you did and count the sentences in future tense. That ratio is the most honest productivity metric most founders will ever run, and it is closely related to why so many side projects fail despite real hours going in.
Test 5: The checklist
Finally the requirement is broken into two to five concrete sub-requirements, each marked met, partial, or missing, each with a note under twelve words that either quotes the exact evidence or names exactly what is absent.
The word limit is not a formatting preference. Short notes cannot hide behind vagueness. "Could be stronger" does not fit the standard; "no reply from any of the four" does.
How the grade is actually calculated
Only now does a number get assigned, and two hard caps sit above the entire scale:
- If any red flag is present, the grade cannot exceed 3.
- If evidence strength is "none", it cannot exceed 3. If it is "weak", it cannot exceed 6.
Those caps are the point. They make it structurally impossible to score well on unproven work, no matter how much effort went in or how good the write-up is. Within them, the bands are:
- 1–3 — did not really do it. No real evidence, pure future tense, unrelated, or a restatement of the task.
- 4–6 — partially done or done weakly. Some real work, but a meaningful gap, thin evidence, or an adaptation whose reasoning is not yet verified.
- 7–8 — genuinely done to a real standard, with evidence that proves it.
- 9–10 — done excellently, with strong undeniable proof and obvious effort.
Notice where honest, incomplete work lands. Real-but-unfinished work belongs at 4 to 6, not at 1. A rubric that scores everything imperfect as a failure teaches you nothing, and founders stop listening to it.
Grade your own work this week
You do not need our software to run this. Take the last thing you called "done" and go through it in order.

- Write down what the task required and what proof would settle it. Do this before you look at your output.
- Quote the actual evidence you have. If you cannot quote something specific, write "none" and be honest about it.
- Rate that evidence strong, weak, or none.
- Name the single biggest gap in one sentence.
- Check for the three red flags.
- Apply the caps, then score it.
Then answer the question the grader finishes on: what is the one highest-impact action that would make this a 10? One sentence, specific to what you actually produced, never generic.
Most founders discover the same thing on their first pass. The work was real. The proof was not. That is a much easier problem to fix than a bad idea, and it is the gap between activity and what actually counts as validation.
Paul Graham's Do Things That Don't Scale makes a related point from the other direction: the early work that matters is usually manual, unglamorous, and leaves a trail of real conversations with real people. That trail is exactly what this rubric looks for. Y Combinator's essential startup advice lands in the same place — talk to users, build something they want, and measure it honestly.
The bottom line
Effort is not the metric. Evidence is.
The founders who ship are rarely the ones who worked the most hours. They are the ones who, at the end of every week, can point at something a stranger could verify. If you cannot do that, the week did not happen, however busy it felt.
If you want the rubric applied to your work by something that will not let you off the hook, that is exactly what GRILLR does — it turns your idea into dated tasks and refuses to mark them done without proof.
Key takeaways
- Evidence, not effort, is the metric: a bare link or file name proves nothing on its own.
- Write down what the task required BEFORE looking at your output, or you will redefine done to fit what you made.
- Three red flags cap any submission at 3/10: pure future tense, restating the task, and unrelated work.
- Evidence strength caps the score: none caps at 3, weak caps at 6, no matter how much effort went in.
- A reasoned adaptation is not off-topic work; deviating with a concrete reason gets graded on its own merits.
Done reading? Stop planning and start building.
Start building