[{"data":1,"prerenderedAt":344},["ShallowReactive",2],{"blog:\u002Fblog\u002Fgemini-3-7-multiple-witnesses":3,"site":332},{"id":4,"title":5,"author":6,"body":7,"date":315,"description":316,"extension":317,"heroAlt":318,"heroImage":319,"meta":320,"navigation":321,"ogImage":319,"path":322,"seo":323,"stem":324,"tags":325,"updated":330,"__hash__":331},"blog\u002Fblog\u002Fgemini-3-7-multiple-witnesses.md","One lesson, multiple witnesses: why Gemini 3.7 did not become BitterClip's default","Michael Ruescher, Founder",{"type":8,"value":9,"toc":303},"minimark",[10,14,17,24,31,36,39,42,61,64,67,71,83,86,97,100,103,107,110,181,184,192,195,199,202,205,208,211,214,226,229,233,236,239,245,260,263,266,269,273,279,282,285,289,292,295,300],[11,12,13],"p",{},"Gemini 3.7 looked like an easy upgrade. In our offline comparison its\nprovider-attempt latency was 47.4% lower, and its outputs cost an estimated 20%\nless to process. Then it failed\nthe test that mattered most: running through BitterClip's real production job,\nit placed five observations outside the video chunk it was describing. The same\nresponse also contained an interval that ended before it began.",[11,15,16],{},"We did not make it the default.",[11,18,19],{},[20,21],"img",{"alt":22,"src":23},"Gemini 3.7 had lower latency and estimated cost in the offline comparison, but failed BitterClip's production temporal-grounding gate.","\u002Fimages\u002Fblog\u002Fgemini-3-7-benchmark\u002Fgemini-3-7-benchmark-chart.svg",[11,25,26,27],{},"That is not the same as saying Gemini 3.7 is a worse model. It is the narrower,\nmore useful conclusion: ",[28,29,30],"strong",{},"Gemini 3.7 did not earn control of BitterClip's source\nclock.",[32,33,35],"h2",{"id":34},"why-one-lesson-needed-two-cameras","Why one lesson needed two cameras",[11,37,38],{},"The test footage was one real outdoor coaching session, recorded at the same\ntime from a static phone and a close first-person camera. One camera could see\nthe shape of the lesson. The other could see details the wide shot lost. In\nsome moments, a person was mostly hidden in one view and clear in the other.",[11,40,41],{},"This is the everyday multi-camera problem in miniature. If a coach gives an\ninstruction, demonstrates it, watches the learner try it, and then offers a\ncorrection, a useful system has to keep several things straight:",[43,44,45,49,52,55,58],"ul",{},[46,47,48],"li",{},"who is teaching and who is responding;",[46,50,51],{},"what each camera actually shows;",[46,53,54],{},"whether an instruction is followed by a visible action;",[46,56,57],{},"which details become knowable only after combining views;",[46,59,60],{},"when the picture is too occluded or ambiguous to support a claim.",[11,62,63],{},"We synchronized the recordings from their audio before asking a model anything.\nThe overlap was fixed by BitterClip, not guessed by Gemini. That distinction is\nimportant: two cameras only become two witnesses after the system establishes\nthat they are showing the same moment.",[11,65,66],{},"We also kept observation separate from editing. A camera can contain useful\nevidence without being the shot an editor should choose. Coverage is not the\ncut.",[32,68,70],{"id":69},"the-benchmark","The benchmark",[11,72,73,74,78,79,82],{},"We tested Google's stable ",[75,76,77],"code",{},"gemini-3.7-flash"," against the exact\n",[75,80,81],{},"gemini-3.6-flash"," baseline. We selected 15 diagnostic moments and 23 individual\ncamera excerpts. They\ncovered setup, demonstration, attempts, corrections, fine body position,\nsustained movement, transitions, occlusion, and deliberately bad pairings. The\nlocked visual comparison used checksummed clips with audio removed, so a model\ncould not smuggle spoken instructions into what we called visual understanding.",[11,84,85],{},"For the cleanest model comparison, Gemini 3.6 and 3.7 received the same media,\nprompt, schema, and transform. Each source-local case ran three times, producing\n69 matched pairs. We recorded the exact requested and served model, latency,\ntoken use, estimated cost, normalized output, and every normalization action.",[11,87,88,89,96],{},"(",[90,91,95],"a",{"href":92,"rel":93},"https:\u002F\u002Fai.google.dev\u002Fgemini-api\u002Fdocs\u002Fmodels\u002Fgemini-3.7-flash",[94],"nofollow","Gemini 3.7 Flash model details",")",[11,98,99],{},"The wider workshop also explored sampled frames, native video, several media\nresolutions and thinking levels, transcript and vision combinations, one-pass\nversus observe-then-synthesize analysis, two videos in one request, and\nduplicate, shifted, mismatched, and withheld-camera controls.",[11,101,102],{},"Not every exploratory arm became a scored result. Thinking-level, free-form,\nand modality comparisons stayed in discovery. A whole-recording response failed\nour deterministic time-bound checks, so it produced no efficacy result. We also\ndid not finish a blinded comparative human score after the preregistered\nrepeatability gate failed. Guided human spot checks improved the truth set, but\nthey are not a substitute for that comparison.",[32,104,106],{"id":105},"what-looked-better","What looked better",[11,108,109],{},"All 138 requests completed—69 per model—forming 69 matched pairs. On those\npairs, Gemini 3.7 was clearly better operationally:",[111,112,113,133],"table",{},[114,115,116],"thead",{},[117,118,119,123,127,130],"tr",{},[120,121,122],"th",{},"Measure",[120,124,126],{"align":125},"right","Gemini 3.6",[120,128,129],{"align":125},"Gemini 3.7",[120,131,132],{"align":125},"Difference",[134,135,136,153,167],"tbody",{},[117,137,138,142,145,148],{},[139,140,141],"td",{},"Mean paired provider-attempt latency",[139,143,144],{"align":125},"baseline",[139,146,147],{"align":125},"lower",[139,149,150],{"align":125},[28,151,152],{},"47.4% lower",[117,154,155,158,160,162],{},[139,156,157],{},"Mean paired estimated provider cost",[139,159,144],{"align":125},[139,161,147],{"align":125},[139,163,164],{"align":125},[28,165,166],{},"19.9% lower",[117,168,169,172,175,178],{},[139,170,171],{},"Lexical claim-set repeatability",[139,173,174],{"align":125},"0.333",[139,176,177],{"align":125},"0.341",[139,179,180],{"align":125},"slightly higher",[11,182,183],{},"The latency result is a mean paired difference. Its window-clustered 95%\ninterval was 42.9% to 52.0% lower. The mean paired estimated-cost difference had\na 95% interval of 15.3% to 24.3% lower. Separately, aggregate estimated provider\ncost per analyzed minute was $0.0649 for 3.6 and $0.0513 for 3.7, a 21.0% lower\naggregate ratio.",[11,185,186,187,96],{},"Gemini 3.7 was not on a cheaper tariff. Google listed 3.6 and 3.7 at the same\ntemporary promotional Standard rates through December 31, 2026. The difference\ncame from the observed token and cache mix in these runs. These are rate-card\nestimates, not a reconciled provider invoice. They also measure provider\nattempts, not synchronization, video preparation, queues, storage, or the full\nBitterClip job. (",[90,188,191],{"href":189,"rel":190},"https:\u002F\u002Fai.google.dev\u002Fgemini-api\u002Fdocs\u002Fpricing",[94],"Google's current pricing",[11,193,194],{},"The model-only numbers gave us a reason to continue. They did not give us a\nreason to switch production.",[32,196,198],{"id":197},"the-multi-camera-result-was-not-stable-enough","The multi-camera result was not stable enough",[11,200,201],{},"We registered an intentionally strict median Jaccard threshold of 0.85 before\nseeing any provider output. The exact lexical projection used to calculate that\nscore was locked later: after repeat-2 execution had begun, but before semantic\ninspection and before repeat 3. For the same system, moment, and witness, it\nnormalized the set of claims and measured how much the three runs overlapped.",[11,203,204],{},"The observed median across the wider confirmation suite was 0.320.",[11,206,207],{},"Gemini 3.7's source-local slice scored 0.341, slightly above Gemini 3.6 at\n0.333. That tiny difference does not show that 3.7 was worse. It shows that\nneither model came close to our absolute rule. The metric is lexical, so it can\npunish harmless paraphrase, and 0.85 may have been too ambitious. But changing a\nrule after seeing the result would turn the benchmark into a negotiation.",[11,209,210],{},"The poison tests supplied a second warning. In two of seven registered units,\nthe structural presence or absence of a cross-witness relationship array changed\nacross repeats. These were not human-adjudicated false same-moment claims. A\nseparate duplicate-media control showed why the architecture matters: label the\nsame bytes as two witnesses and a model may still describe their agreement as\ncorroboration.",[11,212,213],{},"The safer design is now clear:",[215,216,217,220,223],"ol",{},[46,218,219],{},"keep each camera's observations intact;",[46,221,222],{},"establish media identity, synchronization, and overlap outside the model;",[46,224,225],{},"let a later synthesis cite those observations without rewriting them.",[11,227,228],{},"When synchronization is missing, BitterClip should forbid a combined claim. A\nmodel obeying that rule has not detected a mismatch; it has followed a guard the\nhost already proved.",[32,230,232],{"id":231},"then-the-production-canary-failed","Then the production canary failed",[11,234,235],{},"The offline benchmark used a candidate inline-media adapter. BitterClip's real\nvisual-analysis job is different: it downloads the Recording, creates silent\nchunks, uploads them through Google's Files API, waits for them to become active,\nasks for structured observations, normalizes local times, stitches the chunks,\nand writes source-linked understanding.",[11,237,238],{},"So we added a narrow canary path before considering a default-model change. The\ncanary was deliberately single-source: it tested the current B5 production job,\nnot multi-camera synthesis. It required one exact model for every chunk,\npinned requested and fallback to the same exact ID to prevent cross-model\nfallback, refused normalization repairs, prevented the candidate from entering\nproduct understanding, and required remote-file cleanup.",[11,240,241,242,244],{},"One production canary ran against a real Recording. The response to the canary\nrequest pinned to ",[75,243,77],{}," contained:",[43,246,247,254,257],{},[46,248,249,250,253],{},"five visual observations outside the declared ",[75,251,252],{},"0.000–194.321"," second chunk;",[46,255,256],{},"one reversed visual-observation interval;",[46,258,259],{},"one boundary cue outside the same chunk.",[11,261,262],{},"BitterClip's strict normalizer rejected the result after 431.8 seconds. No\ncandidate annotations entered product understanding. The saved annotation\nprojection remained byte-for-byte identical, the source and Project settings\nremained off, and the temporary Gemini Files store was empty after cleanup.",[11,264,265],{},"This is the strong reason not to upgrade yet. BitterClip treats source media as\ntiming authority. If an observation points outside the material it claims to\ndescribe, every downstream use becomes suspect: grounded questions, visual\nannotations, highlight discovery, and eventually camera decisions. Faster wrong\ncoordinates are not a production improvement.",[11,267,268],{},"It exposed a failure the successful development example did not; one run in\neither direction cannot estimate the failure rate.",[32,270,272],{"id":271},"what-shipped-and-what-did-not","What shipped, and what did not",[11,274,275,276,278],{},"The default model is still ",[75,277,81],{},". We did not need a rollback because\n3.7 never became the default.",[11,280,281],{},"The safety work did ship. BitterClip can now run an entitled, non-promoting\nmodel canary through the real job path; assert the exact served model for every\nchunk; pin requested and fallback to the same exact ID to prevent cross-model\nfallback; fail on timestamp repairs, drops, or clamps; isolate candidate output\nfrom normal recovery and product indexes; and record cleanup without persisting\nprovider file handles.",[11,283,284],{},"The offline workshop's rate-card estimate was $10.46, including $3.50 in\nconservative reserves for calls without returned usage. Development job-path\nchecks added about $0.28. The failed production canary returned no usable usage\nreceipt, so we reserve another $0.50 rather than invent a smaller bill. Total\nconservative exposure remained under $11.25 against a $50 cap.",[32,286,288],{"id":287},"what-would-change-the-decision","What would change the decision",[11,290,291],{},"We will test 3.7 again, but not by repeatedly pulling the canary lever until one\nrun happens to pass. The next attempt needs a declared prompt or schema change\nfor chunk-local time, a newly registered series of production canaries, and the\nsame fail-closed checks. Multi-camera synthesis still needs a durable blinded\nhuman review after the deterministic gates pass.",[11,293,294],{},"So the result is neither \"Gemini 3.7 is bad\" nor \"benchmarks do not matter.\"\nThe result is more concrete:",[11,296,297],{},[28,298,299],{},"Gemini 3.7 did not earn the source-scoped production default because the exact\ncanary broke the source-time contract. Separately, the proposed multi-witness\narchitecture failed its repeatability and poison-stability gates.",[11,301,302],{},"Speed bought the model another test. It did not buy it the clock.",{"title":304,"searchDepth":305,"depth":305,"links":306},"",3,[307,309,310,311,312,313,314],{"id":34,"depth":308,"text":35},2,{"id":69,"depth":308,"text":70},{"id":105,"depth":308,"text":106},{"id":197,"depth":308,"text":198},{"id":231,"depth":308,"text":232},{"id":271,"depth":308,"text":272},{"id":287,"depth":308,"text":288},"2026-08-14","Gemini 3.7 had 47.4% lower provider-attempt latency and 20% lower estimated cost offline. Then a production canary returned invalid source times. Here is why BitterClip stayed on 3.6.","md","A benchmark card reading Lower latency offline, rejected by the production gate, with Gemini 3.7 latency, cost, and temporal-gate results.","\u002Fimages\u002Fblog\u002Fgemini-3-7-benchmark\u002Fgemini-3-7-benchmark-og.png",{},true,"\u002Fblog\u002Fgemini-3-7-multiple-witnesses",{"title":5,"description":316},"blog\u002Fgemini-3-7-multiple-witnesses",[326,327,328,329],"Gemini","visual understanding","benchmarks","engineering",null,"bVJ3diLWdobfLluS2Nyv51Hr7_TkCoyX0V-4CknBu20",{"id":333,"app_origin":334,"extension":335,"mcp_resource_url":336,"meta":337,"og_image_default":338,"pricing_url":339,"signup_url":340,"stem":341,"support_email":342,"__hash__":343},"site\u002F_data\u002Fsite.yml","https:\u002F\u002Fapp.bitterclip.com","yml","https:\u002F\u002Fapp.bitterclip.com\u002Fmcp",{},"\u002Fimages\u002Fbitterclip-og.png","https:\u002F\u002Fbitterclip.com\u002F#pricing","https:\u002F\u002Fapp.bitterclip.com\u002Fsign_up","_data\u002Fsite","hello@bitterclip.com","Tgsbw67aXUlzV7fI_vtqnk0QtrKrpPlvl81_t8jnOlY",1786738128208]