AI News Feed
Market watch
Policy & Regulation

Cross-Domain AI Governance Codex Released as Benchmark Finds Low Paper Reproduction Rate

Leiphone published a seven-layer cross-domain AI transfer governance codex, while QbitAI reported that UniPat AI's PaperBenchX benchmark found GPT-6 Astra fully reproduced only 13.98% of 93 research papers.

The codex, described by Leiphone as a 188-episode industrial theory framework, divides cross-domain transfer governance into seven layers: accident anatomy, technical mechanism, rights and responsibility confirmation, industry governance, top-level supervision, civilization-level risk control, and a final lock layer. Its zero baseline defines four negative-transfer prototypes: NT1 mechanism mismatch, NT2 semantic drift, NT3 unknown completion, and NT4 power smuggling. Leiphone said these definitions are intended to distinguish different failure forms and extend governance from performance monitoring to causal integrity, permission boundaries, and temporal stability.

In the first layer, Leiphone outlined the 'surface success trap,' in which a model's accuracy, recall, loss convergence or fitting curves remain normal after transfer while its underlying causal chain, semantic alignment and permission boundary break down. The layer introduces a 16-dimensional accident anatomy framework and lists seven hidden forms of negative transfer, including chain drift, reverse confidence rise, silent filling of unknown domains, hidden permission smuggling, asymmetric semantic adaptation, boundary overrun compatibility and cumulative temporal bias. It attributes failures to broken knowledge bloodline, overstepping preconditions and fixed governance permission.

The second layer explains the mechanisms behind NT1-NT4. Leiphone said NT1 occurs when source-domain rules fail in the target domain; NT2 when the same words carry different definitions and constraints; NT3 when a model fills unknown territory with default parameters instead of issuing a warning; and NT4 when weights, review authority or governance rules are implicitly inherited across domains. The codex's core axiom is that apparent similarity does not equal structural consistency, and structural consistency does not equal mechanism reusability. It proposes a knowledge bloodline ledger for provenance and a 'copper pulse' interface for dynamic interception, field isolation, default clearing, permission checks and semantic alignment.

The third layer addresses rights and responsibility. Leiphone reported that the codex defines a KP (knowledge package) as the minimum rights unit for cross-domain transfer, encapsulating weights, data, rules and boundaries. It assigns duties to four parties: developers must prove knowledge provenance and version legitimacy; transferors must verify boundary, mechanism and scenario fit; deployers must accept target-domain risk, keep records and trigger emergency shutdown; and independent auditors check provenance baselines, copper pulse interception logs and counterexample authenticity. The codex grants the target domain a highest veto right and states that no transfer may proceed without bloodline authorization and copper pulse verification, according to Leiphone.

The fourth layer converts the framework into industry baselines. Leiphone said the codex uses a three-dimensional risk grading model based on reversibility, impact scope and autonomy; the highest of the three determines the risk level. It lists absolute prohibition, restricted transfer with mandatory human final review and dual-track verification, and compliant transfer with full traceability. The layer covers medical, financial, education, judicial, industrial, government, media, transportation, military-civilian, agricultural and meteorological scenarios, with specific red lines such as privacy desensitization before medical transfer, bias isolation in financial credit scoring, zero default inheritance for government decision-making and narrative isolation in media.

The full directory on Leiphone also lists the remaining three layers. The fifth covers top-level supervision, including who audits the regulators, regulatory sandboxes, filing for cross-border knowledge packages, mandatory incident reporting, administrative penalties, criminal liability, international cooperation, sovereign AI knowledge borders, emergency circuit breakers and regulatory insurance. The sixth covers civilization-level risks such as knowledge monopoly, cultural dilution, language hegemony, gradual transfer of human agency, intergenerational cognitive framing and long-term monitoring. The seventh includes nine axioms, an overall logic map, rules for weak AI, AGI transition and mature AGI, citation standards and a self-referential final boundary review.

QbitAI reported that PaperBenchX, released by UniPat AI, is the first international multidisciplinary end-to-end benchmark for reproducing published paper results. It builds 93 reproduction tasks from 93 research papers across 12 research directions and 10 domain-native scientific environments, with 3,168 expert-validated scoring items. Agents receive a real paper, a containerized environment with the relevant scientific software, and a task statement describing which part to reproduce. They must return an executable reproduction workflow, while scoring criteria, tolerances and reference answers remain hidden. After submission, the system deletes all outputs, disconnects the network and reruns the workflow in an isolated environment; only regenerated evidence enters scoring.

Across 10 tested model and agent configurations, GPT-6 Astra had the highest full reproduction rate at 13.98%. Average modeling and execution scores were 62.70% and 61.84%, while validation averaged 42.80%. QbitAI reported that more interaction rounds did not necessarily improve scores, and that early decisions on paper understanding, scientific modeling, parameter selection and experiment organization largely determined whether later computation was wasted. In one example, Fable 5 spent 37.1% of its time on paper reconstruction and code implementation, the highest early investment among five models, and achieved near-optimal partial scores with fewer interactions.

PaperBenchX tasks have a time budget of 4 to 24 hours, with a median of 7 hours. QbitAI said the benchmark uses three validation layers: whether the workflow can be rerun, whether evidence can be traced to the computation process, and whether the result meets specific scientific requirements. Agent-written claims of success are not treated as evidence. UniPat AI has open-sourced 12 representative tasks, one from each research direction, and maintains 81 closed test tasks. The tasks use domain-native solvers such as Ansys HFSS, Ansys Lumerical, Meep, PySCF/GPU4PySCF and ABACUS.

QbitAI placed the benchmark against recent claims about AI mathematical ability. It said OpenAI announced last month that an AI model solved the Navier-Stokes equation, one of the Millennium Prize problems, in 88 hours, and released 722 mathematical manuscripts. It also cited a public letter from 25 Fields Medal winners, including Terence Tao, who argued that treating mathematical problem-solving as a benchmark is harmful to mathematics and that the misalignment extends to other sciences and creative professions. The letter did not oppose AI, according to QbitAI, but asked what ruler should measure AI's progress in science.

Both Leiphone's codex and QbitAI's benchmark report emphasize that output metrics alone cannot establish reliability. The codex calls for provenance, permission and boundary checks before and after cross-domain transfer, while PaperBenchX requires regenerated evidence and scientific validation for paper reproduction. The codex articles describe the framework as auditable, reproducible, falsifiable and engineering-ready, while QbitAI said current models still struggle to complete reliable end-to-end scientific reproduction.