Run #775
done
rubric v1
session 79416979
gemini-3.7-flash
2026-09-21 02:21
Review team
no_issue
Product AI
nothing found
teaching Yes · mistakes No · satisfactory Yes
Each verdict takes your own. Yours is stored against the stack's, so “where is it wrong most often” is a question with an answer.
-
seen · 100% confidentWork is the right way up The video feed is upside down.Across frames 2–7, all text (e.g. 'Acid -> PH value' and Hindi handwriting) is inverted at 180 degrees orientation.
-
Your view
-
seen · 90% confidentWriting is legible at video size Handwriting is clear and legible despite orientation.In frames 6 and 7, text such as 'Acid -> PH value (0-7)', 'strong acid, weak acid', and 'Blue litmus paper -> Red litmus paper' is clearly readable.
-
Your view
-
seen · 95% confidentThe work fills the frame Notebook page fills most of the frame.In frames 2–7, the spiral notepad is centered and occupies the vast majority of the camera view.
-
Your view
-
seen · 95% confidentNothing distracting in shot Workspace is clean with minimal surroundings visible.Frames 2–7 show only the notebook, the tutor's hand, a pen, and the dark desk surface on the borders.
-
Your view
-
seen · 80% confidentThe student's problem is established No clear problem statement or original question is displayed.Frames 2–7 show notes being taken about acids and pH, but the original question or prompt being answered is not visible.
-
Your view
-
seen · 90% confidentWork proceeds in followable steps Work progresses logically across frames.Frames 3 through 7 show sequential additions: first the pH range, then the pH scale line, and finally the litmus test effect.
-
Your view
-
seen · 85% confidentIt reaches an answer The explanatory notes are completed.By frame 7, the explanation covering acid pH range, scale markings, and litmus paper reaction is fully written out.
-
Your view
-
measuredSomething was actually written Not applicable to a camera session
-
heardThe concept is named before the working Not assessedOnly 157 characters of speech after cleaning (200 were the language-detection artefact) — too little to judge teaching from.
-
Your view
-
heardExplained rather than jumped to the answer Not assessedOnly 157 characters of speech after cleaning (200 were the language-detection artefact) — too little to judge teaching from.
-
Your view
-
heardSolved step by step, not in one leap Not assessedOnly 157 characters of speech after cleaning (200 were the language-detection artefact) — too little to judge teaching from.
-
Your view
-
heardThe session opens properly Not assessedOnly 157 characters of speech after cleaning (200 were the language-detection artefact) — too little to judge teaching from.
-
Your view
-
measuredNo long dead air 0 gaps over 6.0s, 0s total (0% of the session), longest 0sSilence is below -35.0dB — a crude measure on a phone microphone, so treat a low count as weak evidence rather than proof the tutor was talking.
-
Your view
Observations judged, but not counted as the lab objecting
-
measuredThe camera holds still picture moves 37.6 on average between keyframes, 0.6 hard jumps/minConsistent with the camera being handheld — a stand would remove this entirely. A handheld phone measured 35.1. No steady-camera session has been measured yet, so the lower band is a guess.Not counted: Measured over 183 known-bad and 29 known-good camera sessions (2026-09-21): raised on 87% of the bad and 86% of the good, and the motion distributions are the same shape — median frame difference 25 on the bad set, 27 on the good. No threshold on this number separates them, so it is reported and not counted.
-
Your view
-
measuredExposure is usable average brightness 118.8, 34.0% of frame very dark, 13.0% blown outBelow 70 average, or over 25.0% of the frame too dark to read.Not counted: Same two sets: it raised 6 of 183 known-bad and 2 of 29 known-good sessions. Both too few to mean anything, and the brightness distributions overlap almost exactly. The earlier reading that it 'fires more on clean sessions' was two sessions out of twenty-nine — noise, not an inverted rule.
-
Your view
-
heardUnderstanding of the concept is checked Not assessedOnly 157 characters of speech after cleaning (200 were the language-detection artefact) — too little to judge teaching from.Not counted: Raised on 87% of known-bad and 80% of known-good sessions. The criterion is honest about why in its own `why`: the transcript has no speaker labels, so a real check and a rhetorical 'ok?' are often indistinguishable, and the bar ends up one almost no session clears. Still judged and shown, because reading it beside a recording is how it gets better.
-
Your view