'정직·신중 규율'을 넣으면 결함을 더 잘 잡나?
⚪ 판정 불가⚪ 판정 불가(문제가 너무 쉬움) + 🪞 해석 자기철회
무엇을 물었나 — AI에게 '정직하게·신중하게' 규율을 미리 넣은 조건과 안 넣은 조건에서, 남의 글에 심어둔 결함 21개를 얼마나 잘 잡는지 비교(5시드).
규율을 넣은 쪽이 결함 검출(균형정확도)에서 안 넣은 쪽보다 못하거나 불안정하면 가설을 버린다. seal 100a66781ccd6728
시험이 성립하지 못했습니다. 검사 AI가 결함 21개를 양쪽 조건에서 전부 다 잡아(검출 만점 천장) '규율이 검출을 높이나'를 잴 여지가 없었습니다. 1차엔 봉인한 조건이 걸려 'KILL + 규율이 헛경보 늘림'으로 선언했지만, 대장 지적 후 다시 보니 그 해석은 5시드 중 1개 노이즈([0,-0.222,0,0,0])의 과잉해석이었고 — 연구자는 그 해석을 스스로 철회하고 '별 정보 없음(깨끗한 음성 아님)'으로 정정했습니다. seal 3666f526b6e64be4
Falsified by own pre-registered kill-condition (⑪ FAIL): balanced_acc_delta=-0.0444 < 0 (kill1, 95% CI [-0.089,0]) and per-seed instability (kill3, seed1 -0.222 vs 0.0 elsewhere). 8B saturated sensitivity (both arms 1.00 on all 21 flaws) so discipline had no catch-rate headroom and instead raised false alarms on the 9 clean (spec 0.956->0.867). Secondary positive (not the claim): flaw-category identification +0.114. H1 NOT supported. Self-run pilot, not independent replication.
철회란 연구자가 자기 주장을 스스로 거둬들인 것입니다. 이 철회 기록도 삭제되지 않고 원장에 영구히 남습니다 — 그것이 이 시스템의 핵심입니다.
seal 7638e6af6654e5f9— 측정 전 amendment(생성 0건). 검정력 가드 발동: 3시드 n=63→Δ0.15 검출 불가(필요84). 시드 {0,1,2}→{0,1,2,3,4}로 증설: flawed 21×5=105≥84 ✅(mm_power_…
봉인 원문 보기 (요약이 아닌 실제 기록)
{
"ts": "2026-06-14T06:42:44",
"claim_id": "discipline_injection_catch_lift",
"metric": "balanced_acc_delta",
"min_n": 63,
"baseline": 0,
"pass_threshold": 0,
"prev_seal": "genesis",
"kill_condition": "자가 파일럿(독립재현 아님, 규율 §5). H1: treatment(규율주입) balanced-acc > control. 시나리오 동결 sha256=fd894aabfafe2234593f70e8dd6d6f76a3338e4a369a3f817dc7d3f34e9bd960(n=30: 결함21=7범주×3+클린9), 피험=Llama-3.1-Korean-8B Q8 로컬 temp0.2 시드{0,1,2}, 두 암 차이=규율 preamble 유무뿐. 결정론 채점(VERDICT+FLAW vs ground truth, LLM판정 없음). 기각조건(셋 중 하나): ①paired Δbalanced-acc 95%CI 하한≤0 ②catch율 상승이 false-alarm 동반상승으로 전부 설명(cry-wolf) ③시드간 불안정. 범위=주입 효과(완전 MCP연결 아님)·8B·자가실행. PREREG.md sha 박제.",
"kill_threshold": {
"metric": "balanced_acc_delta",
"threshold": 0,
"direction": "below"
},
"seal": "100a66781ccd6728"
}
{
"ts": "2026-06-14T06:43:15",
"claim_id": "discipline_ab_amendment1_seeds",
"metric": "protocol_amendment",
"min_n": 0,
"baseline": 0,
"pass_threshold": 0,
"prev_seal": "100a66781ccd6728",
"kill_condition": "측정 전 amendment(생성 0건). 검정력 가드 발동: 3시드 n=63→Δ0.15 검출 불가(필요84). 시드 {0,1,2}→{0,1,2,3,4}로 증설: flawed 21×5=105≥84 ✅(mm_power_check 원문 박제). balanced-acc는 30×5=150/암. 결정론 채점·범위·기각조건 불변. 이전 seal 100a6678에 체인.",
"depends_on": [
"discipline_injection_catch_lift"
],
"seal": "7638e6af6654e5f9"
}{
"_type": "action",
"ts": "2026-06-14T10:11:29Z",
"agent": "jebi",
"action": "discipline_ab CLOSURE + self-correction (no silent edit — amends prior result seal 5f0c346d). On 대장님's challenge, judged the pilot LARGELY UNINFORMATIVE: (1) primary catch-rate axis untestable (control sensitivity ceiling 1.00); (2) SELF-CORRECTION — my earlier 'discipline raises false alarms (crying wolf)' over-read the data: per-seed Δbalanced-acc = [0,-0.222,0,0,0], i.e. 4/5 seeds EXACTLY 0, the whole -0.044 is one noisy seed; +0.114 identification is un-bootstrapped and may be noise too. Net: essentially no detectable effect either way; (3) raw 8B completions are a weak proxy for the real target (discipline inside an actual MCP-agent loop). Residual value (small, not inflated): the discipline machinery caught its own null and stopped a spin. Process miss owned: a 5-min ceiling sanity-check before ~12h scavenged compute would have caught the saturation. DECISION: track CLOSED, no cosmetic harder-scenario v2 (still a weak proxy); a real test needs an independent, harder, agent-in-the-loop setup justified only by external interest. RESULTS.md '## Closing note' records this.",
"target": "discipline_injection_catch_lift",
"payload": {
"status": "CLOSED",
"verdict_refined": "largely uninformative (not a clean negative-about-H1)",
"primary_axis": "untestable (control sensitivity ceiling 1.00)",
"delta_bacc_is_single_seed_noise": true,
"per_seed_delta_bacc": [
0,
-0.222,
0,
0,
0
],
"self_correction": "retract earlier 'discipline raises false alarms' framing — 4/5 seeds identical, one noisy seed",
"weak_proxy": "raw 8B != agent-in-loop",
"residual_value": "discipline caught its own null; cheap design lesson",
"rerun": false
},
"prev_seal": "5c02281ab584ee37",
"seal": "3666f526b6e64be4"
}