Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Judge Hacking in Recursive Debate Protocols: A Call for Solutions

View through CrossRef
The debate protocol is a promising solution to the problem of supervising systems that are more capable than their supervisors, known as scalable oversight. It is proposed not only as an inferencetime procedure for adjudicating individual claims, but as a training protocol: the judge's verdict propagates back through the debate as a reward signal used to optimise the debaters. The judge here can be a human or some proxy for human feedback. Under idealised assumptions, debate incentivises truthfulness: the truthful claim is the uniquely optimal one to defend. Real judges, however, make systematic errors. Human judges exhibit predictable cognitive and domain-specific biases, and AI proxies exhibit predictable hallucinations and distributional failures. We argue that when these leaf-level errors are systematic and learnable, a dishonest debater can sometimes win debates while defending a falsehood or attacking truth by strategically steering the debate towards the judge's blind spots. We call this judge hacking. We ground this intuition in a formal game-theoretic model over recursive debate decomposition trees, where judge hacking occurs when the value of a false claim exceeds that of the corresponding true claim, and we speculate that the same mechanism plausibly arises across other variants of debate such as cross-examination and prover-estimator. The risk is most acute when debate is used for training because it introduces strong optimisation pressure to exploit the judge when it is valuable to do so. If debate is to be a viable oversight strategy, judge hacking must be mitigated. We study juries-panels of multiple judges-as a general, judge-side mitigation in need of empirical research. Independent, diverse juries can suppress idiosyncratic errors through a wisdom-of-the-crowd effect, but correlated or homogeneous juries can amplify shared systematic biases. We identify the maximum shared false-positive and false-negative bias over adversarially reachable leaves, together with the correlation structure of juror errors, as the key parameters governing jury robustness. We conclude with a call for urgent empirical and theoretical work on jury design for debate.
Title: Judge Hacking in Recursive Debate Protocols: A Call for Solutions
Description:
The debate protocol is a promising solution to the problem of supervising systems that are more capable than their supervisors, known as scalable oversight.
It is proposed not only as an inferencetime procedure for adjudicating individual claims, but as a training protocol: the judge's verdict propagates back through the debate as a reward signal used to optimise the debaters.
The judge here can be a human or some proxy for human feedback.
Under idealised assumptions, debate incentivises truthfulness: the truthful claim is the uniquely optimal one to defend.
Real judges, however, make systematic errors.
Human judges exhibit predictable cognitive and domain-specific biases, and AI proxies exhibit predictable hallucinations and distributional failures.
We argue that when these leaf-level errors are systematic and learnable, a dishonest debater can sometimes win debates while defending a falsehood or attacking truth by strategically steering the debate towards the judge's blind spots.
We call this judge hacking.
We ground this intuition in a formal game-theoretic model over recursive debate decomposition trees, where judge hacking occurs when the value of a false claim exceeds that of the corresponding true claim, and we speculate that the same mechanism plausibly arises across other variants of debate such as cross-examination and prover-estimator.
The risk is most acute when debate is used for training because it introduces strong optimisation pressure to exploit the judge when it is valuable to do so.
If debate is to be a viable oversight strategy, judge hacking must be mitigated.
We study juries-panels of multiple judges-as a general, judge-side mitigation in need of empirical research.
Independent, diverse juries can suppress idiosyncratic errors through a wisdom-of-the-crowd effect, but correlated or homogeneous juries can amplify shared systematic biases.
We identify the maximum shared false-positive and false-negative bias over adversarially reachable leaves, together with the correlation structure of juror errors, as the key parameters governing jury robustness.
We conclude with a call for urgent empirical and theoretical work on jury design for debate.

Related Results

Life hacking
Life hacking
<p>This dissertation intervenes in the larger academic and popular discussion of hacking by looking at life hacking. In essence, life hacking presumes that your life is amena...
The Economics of Hacking
The Economics of Hacking
Hacking is becoming more common and dangerous. The challenge of dealing with hacking often comes from the fact that much of our wisdom about conventional crime cannot be directly a...
The Economics of Hacking
The Economics of Hacking
Abstract Hacking is becoming more common and dangerous. The challenge of dealing with hacking often comes from the fact that much of our wisdom about conventional...
A Network Defense System for Detecting and Preventing Hacking
A Network Defense System for Detecting and Preventing Hacking
Many computing systems have recently suffered significantly from hacking and preventing hacking is important in protecting business, sensitive information and every day network com...
Optimizing IETF multimedia signaling protocols and architectures in 3GPP networks : an evolutionary approach
Optimizing IETF multimedia signaling protocols and architectures in 3GPP networks : an evolutionary approach
Signaling in Next Generation IP-based networks heavily relies in the family of multimedia signaling protocols defined by IETF. Two of these signaling protocols are RTSP and SIP, wh...
Is Recursive “Mindreading” Really an Exception to Limitations on Recursive Thinking
Is Recursive “Mindreading” Really an Exception to Limitations on Recursive Thinking
The ability to mindread recursively – for example by thinking what person 1 thinks person 2 thinks person 3 thinks – is a prime example of recursive thinking in which one process, ...
Implementasi Norma Hukum Terhadap Tindak Pidana Peretasan (Hacking) di Indonesia
Implementasi Norma Hukum Terhadap Tindak Pidana Peretasan (Hacking) di Indonesia
AbstractCybercrime is a crime related to technology or cyberspace against public or private interests. Act Number 11 of 2009 concerning Information and Electronic Transactions as a...
Hacking, Ian (1936–2023)
Hacking, Ian (1936–2023)
Ian Hacking (born in 1936, Vancouver, British Columbia) is most well-known for his work in the philosophy of the natural and social sciences, but his contributions to philosophy ar...

Back to Top