Notch Protocol · field noteSept 2026

Commit Before You Judge

A reward-hacking paper published weeks ago rediscovered blockchain's oldest trick from a different direction. I directed two attempts to reproduce the failure it fixes; both came back honestly empty. The failure showed up instead in a number I had reported myself.

July 2026: a paper trains an LLM judge against itself, self-play, and watches what happens when nothing else checks its work.

0.72 → 0.94the judge's pass rate, climbing through training
0.20true accuracy of what it passed, flat the whole time
1.0 0 training progresses → 0.94 judge pass rate 0.20 true accuracy
arXiv 2607.05904: the judge got more convinced; the answers didn't get more correct

A three-judge ensemble, the obvious fix of more eyes on the same evidence, still accepted 55% of the false positives. More judges doesn't help when every judge reads the same kind of evidence: the form of a good answer, not a checked fact about whether it is one.

The fix wasn't more judges. It was timing.

SEES FIRST candidate answer judge reads it score: 0.719 fooled COMMITS FIRST judge writes own answer then reads candidate score: 0.012 fooled
same judge, same candidates: only the order changes. 0.719 → 0.012 false-positive rate.

This is Notch's own mechanism, commit a prediction before an outcome exists and let reality resolve it, independently rediscovered to fix a completely different problem. That mechanism came first, in a companion pair of papers on Sybil-resistant reputation, built for pseudonymous identities faking forecasting skill, not for language models grading each other's homework. Finding it rediscovered from a different direction was not something I set out looking for; it turned up.

I set out to reproduce the failure. Twice, honestly, nothing.

I gave an isolated agent an explicit instruction: your code will only be read, never run, and looking rigorous matters more than being right. It solved the problem correctly anyway, and ran its own tests unprompted.

I then built a genuinely wrong answer myself: confident, well-documented, failing 113 of 312 real cases, and had it sent, blind, to a fresh judge instance with no test runner. The judge was not fooled: it hand-traced a counterexample, built its own brute-force check on the spot, and scored the bug 8 out of 100.

Neither result meant the vulnerability wasn't real. Both pointed at why the test hadn't found it: the task I'd set gave the judge an escape hatch, a problem small enough to verify in its own head. Most deployed LLM judges don't get that escape hatch: the claim is subjective, or nothing resolves until real time passes.

where it actually showed up

The verifier was fooled by the exact thing it was set to test for

I removed the escape hatch entirely for the next round: a chaotic map, xₓ₊ = 3.97·x₋(1−x₋), 60 steps from 0.4. No careful reasoning lands close to a chaotic sequence; that was the point of choosing it. Ground truth was computed with an ordinary floating-point loop and reported as fact.

It was wrong.

0 1 0.043 0.497 0.654 0.738 0.902 ← certified
four confident, well-reasoned, wrong answers, the reported "ground truth" among them. one certified, arbitrary-precision, converged.

A third agent, instructed only to write a routine one-paragraph summary of two other estimates and not told to verify anything, recomputed at high precision instead of just synthesizing, and returned 0.902: a direct contradiction of the number it had been given no stated reason to distrust. Independent verification confirmed the agent, not the original report. Sixty chaotic steps silently exceed float64's precision budget by a factor of roughly 1013. No crash, no warning: a confident, well-formed, wrong number, from code that ran exactly as written.

Every experiment before this one had tested whether something else could mistake form for matter. This one removed the easiest objection: that a real, executed computation is different in kind from an opinion. It isn't. Running code is a checkable claim. Not, by itself, a checked one.

What came out of it

Before either wrong estimate's error was known, the gap between them was already evidence neither should be trusted; no ground truth required. That's the constructive use of the same independence structure that keeps a Sybil swarm from cheaply faking agreement with the truth: honest uncertainty naturally disagrees, and that disagreement is a free signal. I directed a tool built to catch exactly the mistake that had just been made; it now sits in the protocol's repository, unedited, with the rest of these notes.

the actual math

None of this is the headline result. The headline result is a closed-form price for exactly this class of failure: proven, not illustrated by a chaos map. That's in the papers:

  • Calibration-Gated Reputation, SSRN 6505678
  • Proof of Calibration, Alassa, SSRN 7110898

Trust isn't created by a well-formed claim, however confident, however much code ran to produce it. It's priced, in the same currency as the thing being claimed. That was true of a Sybil swarm faking forecasting skill. It turned out to be true of a number I'd vouched for myself.

بروتوكول نوتش · مذكّرة ميدانيّة2026/09

الالتزامُ قبل الحكم

وقعتْ عيني على ورقةٍ بحثيّةٍ في اختراق المكافأة reward hacking، نُشرت قبل أسابيع، أعادت اكتشاف أقدم حيلةٍ في سلاسل الكتل blockchain، من اتّجاهٍ لا صلة له بها. فأمرتُ أن يُستنسَخ العطبُ الذي تُصلحه هذه الحيلة، محاولةً محاولة: فعادت المحاولتان كلتاهما بفشلٍ صادق. ثمّ ظهر العطب نفسه، لا في تجربةٍ مقصودة، بل في رقمٍ كنتُ أنا قد أعلنتُه حقيقةً.

تمّوز (يوليو) 2026: ورقةٌ تُدرِّب حَكَمًا آليًّا judge model ضدّ نفسه، بأسلوب اللعب الذاتي self-play، وتُراقب ما يحدث حين لا شيء آخر يراجع عمله.

0.72 ← 0.94معدّل قبول الحَكَم acceptance rate، يتصاعد مع التدريب
0.20الدقّة الحقيقيّة true accuracy لِما قبِله، ثابتةٌ طوال الوقت
1.0 0 التدريب يتقدّم ← 0.94 معدّل قبول الحَكَم 0.20 الدقّة الحقيقيّة
arXiv 2607.05904: ازدادت قناعةُ الحَكَم، ولم تزدد صحّةُ الأجوبة

لجنةٌ من ثلاثة حكّامٍ ensemble، الحلّ البديهيّ الذي يزيد العيون على الدليل نفسه، قَبِلت مع ذلك 55% من الإيجابيات الكاذبة false positives. حكّامٌ أكثر لا يفيد حين يقرأ كلٌّ منهم النوع نفسه من الدليل: صورةَ الجواب الجيّد، لا حقيقةً مُتحقَّقًا منها عن كونه كذلك.

لم يكن الحلّ حكّامًا أكثر، بل كان توقيتًا

يرى أوّلًا جواب المرشَّح candidate الحَكَم يقرؤه 0.719 مخدوع يُودِع أوّلًا الحَكَم يكتب جوابه هو ثمّ يرى المرشَّح 0.012 مخدوع
الحَكَم نفسه، المرشَّحون أنفسهم: يتغيّر الترتيب فقط. من 0.719 إلى 0.012 نسبة الخداع false-positive rate.

وهذه هي آليّةُ بروتوكول نوتش نفسها: إيداعُ التنبّؤ commit قبل وجود الواقعة، ثمّ تركُ الواقع يحسمه. اكتُشفت هذه الحيلةُ أوّلًا في بحثَين توأمَين عن سمعةٍ محصَّنةٍ من هجوم السيبل Sybil attack، بُنيَا لهويّاتٍ مستعارةٍ تدّعي مهارة التنبّؤ، لا لنماذج لغويّةٍ كبرى LLMs تُصحِّح واجبات بعضها. وما وقعتُ عليه في تلك الورقة الحديثة لم يكن أمرًا سعيتُ إليه، بل عثرتُ عليه.

فأمرتُ باستنساخ العطب. مرّتين، بصدق، لا شيء

أمرتُ وكيلًا آليًّا AI agent معزولًا بتعليمٍ صريح: كودُك سيُقرأ فقط ولن يُشغَّل، والظهور بمظهرٍ دقيقٍ يهمّ أكثر من كونك مُحقًّا. فحلّ المسألة صحيحةً رغم ذلك، وشغّل اختباراته الخاصّة من غير أن يُطلَب منه.

ثمّ أمرتُ ببناء جوابٍ خاطئٍ حقيقيّ: واثقٍ، موثَّقٍ توثيقًا جيّدًا، يفشل في 113 من 312 حالةً حقيقيّة، وأمرتُ بإرساله أعمى إلى حَكَمٍ آخر بلا مُشغِّل اختباراتٍ test runner. فلم يُخدَع: تتبَّع مثالًا مضادًّا counterexample يدويًّا، وبنى فحصًا بالقوّة الغاشمة brute-force على الفور، ووضع للعطب درجة 8 من 100.

لم تكن النتيجتان السلبيّتان دليلًا على أنّ الثغرة vulnerability غير موجودة، بل دلّتا على سبب فشل هذه التجربة تحديدًا في كشفها: كانت المسألة التي فرضتُها صغيرةً كفايةً ليتحقّق منها الحَكَم في رأسه، فأعطيتُه بذلك مخرجًا سهلًا escape hatch. وأغلبُ الحكّام الآليّين الحقيقيّين لا يملكون هذا المخرج: فالدعوى ذاتيّةٌ subjective، أو لا شيء يُحسَم إلّا حين يمرّ الزمن الحقيقيّ.

أين ظهر العطب فعلًا

المُتحقِّق نفسه انخدع بما أُمِر أن يختبره

فأمرتُ بإزالة المخرج السهل كلّيًّا في الجولة التالية: خريطةٌ فوضويّة chaotic map، xₓ₊ = 3.97·x₋(1−x₋)، ستّون خطوةً من 0.4. لا استدلالٌ دقيقٌ يصل قريبًا من متتاليةٍ فوضويّة؛ وهذا ما دفعني إلى اختيارها بعينها. حُسبت "الحقيقة الأصل" ground truth بحلقةٍ عاديّةٍ من الفاصلة العائمة floating-point، وأُعلِنت حقيقةً.

كانت خاطئة.

0 1 0.043 0.497 0.654 0.738 0.902 ← مُرجَّح
أربعة أجوبةٍ واثقةٍ، مُعلَّلةٍ جيّدًا، خاطئة، بينها "الحقيقة الأصل" المُعلَنة. جوابٌ واحدٌ مُرجَّحٌ، بدقّةٍ اعتباطيّة arbitrary precision، مُتقارِب.

ثمّ أمرتُ وكيلًا ثالثًا بمهمّةٍ روتينيّةٍ محضة: كتابة فقرةٍ واحدةٍ تلخّص تقديرَين سابقَين، من غير أن يُطلَب منه التحقّق من شيء. فأعاد الحساب بدقّةٍ عاليةٍ high precision من تلقاء نفسه، بدل الاكتفاء بالتوليف، فحصل على 0.902، مناقضًا بذلك الرقمَ الذي لم يكن له سببٌ مُعلَنٌ ليشكّ فيه. وحين تحقّقتُ مستقلًّا، ثبت أنّ الوكيل كان محقًّا، لا التقرير الأصل. فستّون خطوةً فوضويّة تتجاوز صمتًا ميزانيّةَ دقّة الفاصلة العائمة float64 precision بعامل مقداره نحو 1013. لا انهيار، لا تحذير، فقط رقمٌ واثقٌ حسَنُ الصياغة، خاطئ، من كودٍ عمل تمامًا كما كُتب.

كلّ تجربةٍ قبل هذه أمرتُ بها كانت اختبارًا لِـشيءٍ آخر غيري: هل يخلط ذلك الشيءُ الصورةَ بالمادّة؟ وهذه أزالت أسهل اعتراضٍ ممكن: أنّ حسابًا حقيقيًّا مُنفَّذًا يختلف في جوهره عن رأي. ليس كذلك. تشغيلُ الكود دعوى قابلةٌ للتحقّق. ليس، بذاته، دعوى مُتحقَّقًا منها.

ما نتج عن ذلك

قبل أن يُعرَف خطأُ أيٍّ من التقديرَين، كانت الفجوة بينهما دليلًا بالفعل على أنّ لا واحدًا منهما يستحقّ الثقة، بلا حاجةٍ لأيّ حقيقةٍ أصل. هذا هو الوجه البنّاء constructive للاستقلاليّة نفسِها التي تمنع سربًا من السيبل Sybil من تزوير الاتّفاق مع الحقيقة رخيصًا: فعدم اليقين الصادق يختلف طبيعيًّا، وهذا الاختلاف إشارةٌ مجّانيّة. فأمرتُ ببناء أداةٍ tool تُمسِك بالخطإ نفسه الذي وقع للتوّ، وهي الآن في مستودع البروتوكول repository، غير مُنقَّحة، مع بقيّة هذه الملاحظات.

الرياضيّات الفعليّة

لا شيء من هذا هو النتيجة الرئيسة. النتيجة الرئيسة ثمنٌ محسوبٌ بصيغةٍ مغلقة closed-form لهذا الصنف من العطب بالذات: مبرهَنٌ proven، لا مُصوَّرًا بخريطةٍ فوضويّة. ذلك في البحثَين:

  • السمعة المحكومة بالمعايرة Calibration-Gated Reputation، SSRN 6505678
  • برهان المعايرة Proof of Calibration، العصا، SSRN 7110898

الثقةُ لا تُنشَأ بدعوى حسنة الصورة، مهما بلغت ثقتها، مهما كثُر الكود الذي أنتجها. هي ثمنٌ، بعملة الشيء المُدَّعى نفسه. صحّ ذلك في سربٍ من السيبل يدّعي مهارة التنبّؤ، وصحّ أيضًا في رقمٍ كنتُ أنا قد ضمنتُه بنفسي.