最新下载
热门教程
- 1
- 2
- 3
- 4
- 5
- 6
- 7
- 8
- 9
- 10
近似自然潜变量具有精确价格 - LessWrong
时间:2026-07-07 08:31:00 编辑:袖梨 来源:一聚教程网
Second content post in a planned cluster on exact results for natural latents.
See the introduction and the previous post. In this post, I share some key theoretical results that this framing generates. This post has more math in it than the last one, most of which I've banished to collapsible sections.[1]
I assume you've read the previous posts, but I tried to make the pedagogical arc make sense even if you start with this one. It helps if you have some familiarity with natural latents, canonical correlation analysis[2], and information theory.
Consider an Elephant in a Room
Say Alice and Bob have access to two different camera feeds observing the same room.

Alice and Bob watch the same room through cameras in opposite corners, but they can only see their own feeds. Their feeds
- redundant: each of them can pin it down from their own feed alone, and
- mediating: given
, the feeds carry no further information about each other; it explains all of their agreement.
"What's in the room?" is the intended example.[3] The last post showed the two conditions correspond to the two classical notions of common information (Gács–Körner and Wyner), that an exact natural latent exists iff
So the natural latents framework runs on approximation, and each condition above gets an associated error term. The mediation error
Let's call that stuck part the remainder. (Symmetrically

Then, an obvious question to ask: how small can the two errors be, together? How natural can natural latents be?
As I noted in the introductory post, this question is hard because the core objects have no generic closed form. This post answers the question in the jointly Gaussian case, where everything does have a closed form.
Claim.For jointly Gaussian views with correlation
If it varied at all, we'd have a feature of Alice's feed and a feature of Bob's that agree with correlation exactly 1, which two views with
Top: No informative function can exactly agree, Bottom: Guaranteed agreement carries no information.
A Distribution-free Sum Rule
The two errors are connected by an identity we can obtain with some simple algebra. The derivation uses the chain rule for mutual information, see below.
Derivation of the Sum Rule
The chain rule for mutual information is
which says: what the pair
Bob's remainder. Expand what the two feeds together say about
Rearranged:
Alice's remainder. Now expand what the pair (concept + Alice's feed) says about Bob's feed, in two different orders:
The middle expression takes Alice's feed first: her feed predicts Bob's by
Now adding, the
So, for any latent over any pair of views, with any[5] distribution:
The interpretation: every bit of the concept that can't be read off one feed alone shows up as either bits the concept carries beyond the shared information, or bits of agreement it leaves unexplained.
This already gives one part of the answer we're looking for. If we demand zero mediation error, then the total remainder is
The views can be almost identical, and the best fully-explanatory concept still can't be pinned down from one of them to better than half a bit.
The Exact Tradeoff Curve
Fix a mediation budget
Then, for each
where
Derivation of the Curve
The latent that sits exactly on the tradeoff curve is the Wyner mediator from the previous post with one change. Take a standard normal
for
As a sanity check, at
Its redundancy errors come out of the sum rule. Everything is jointly Gaussian, so
with the two remainders equal by symmetry.
So this construction achieves the curve, but why does nothing beat it?
Reading the sum rule backwards: at a fixed mediation budget
Importantly, this minimum is over any possible latent. Substituting their closed form back through the sum rule gives a floor of
The derivation follows. Write everything as logs of the four factors
The sum rule at mediation budget
the
We rearrange to get the desired expression:
The attached script checks the achieving latent against the closed form and runs a random search over valid latents; none crosses the floor set by the curve. That search is the gray cloud in the figure below.
And here is that curve in a plot:

Consider the left side of the curve. At exact mediation, the remainder is
The Cliff
The cost of a fully shared concept rises as the views converge, all the way until the instant they coincide, where it falls off a cliff. At

So near-perfect agreement is the most expensiveplace to want a fully shared concept, and only exact agreement makes it free. If you want the free regime, we must make the views exactlysimilar.[7]
The Exchange Rate
The curve is also steepest at the left: the first sliver of mediation tolerance is the most valuable. Tolerate just
At the right end the trade flattens to 8 bits of agreement explained per bit of remainder taken on, at
At
My interpretation is that the existence of a worthwhile shared concept depends on a sharp threshold: high-overlap views share richly and cheaply, views with low-or-middling overlap have nothing that's "worth it" to share.
The Floor
Maybe you don't care about either error separately and just want the best balanced concept, minimizing the worse of the two errors. That's the point where the curve crosses the diagonal:
So, this is a floor under naturality itself. Below
Example: A Biased Die
Of the above, the half-bit cliff in particular is a weird result. It's discontinuous and counterintuitive. Could it be an artifact of Gaussian algebra? Here it is showing up in a familiar example, the biased die.
Consider a die with an unknown bias. Alice gets one long run of rolls and Bob gets another independent run. The bias is the thing they share: it's the entire reason their runs look alike, so it's the exact mediator, and mediation error is zero by construction. The redundancy error is the question: how much of the bias is stuck outside Alice's run?
For estimating the bias, Bob's run doubles Alice's sample. Doubling the sample halves her posterior variance, and halving a variance is worth
Below is this experiment computed in the Beta-binomial model: the remainder is

What this Means
Two observers looking at the same world from different views never hold exactly the same concept. Any concept available to both carries a remainder, a part stuck on the other side that only pooling views would resolve. The size of the remainder is set by how much the vantages overlap, reaching zero only when they coincide exactly.
Recall that the Natural Abstractions agenda is motivated by the hope that a capable AI—modeling the same world as we do—will form the concepts that we form such that we could find ours in it or point at them.
The tradeoff curve says[11] that the AI's version of any shared concept holds a remainder relative to ours, and ours relative to its. The potentially useful question: how large is the remainder for the concepts we care about, and when is it small enough to rely on? [12]
What's Next
Agents track many features of an object at once (e.g. shape, color, position), with each overlapping a different amount. Next post: the optimal shared concept over many features keeps the strongly overlapping ones and discards the weakly overlapping ones. Combined with correlations decaying over distance, this says that shared concepts will drop modes (features) one at a time discretely rather than blurring them gradually. I think this could be tested on learned representations in neural networks.
Also, everything in this post allows the latent to be stochastic. Whether a deterministicconcept can always do essentially as well is Wentworth & Lorell's $500 bounty question; the machinery here settles only the Gaussian case (with linear error transfer), which I plan to write up later in the sequence.[13]
Every number and plot in this post can be generated from this script.
- ^
Conventions: Logarithms are base 2, quantities are in bits. A python script reproducing the numbers shown in this post is linked at the end.
- ^
For those who want to get a deep understanding of this post, I highly recommend watching this presentation introducing CCA and CICA by Prof. Michael Gastpar, whose work I build on.
- ^
See my comment on the last post, for a more detailed breakdown of what each classical object corresponds to in the case of the diagram.
- ^
The claim holds for any number of views, and for the stronger "recoverable from each complement-of-a-view" variant of redundancy; that version needs one extra idea (the latent becomes an invariant of a Gibbs resampling chain, and ergodicity forces invariants to be constant).
- ^
The MI algebra does not assume any type of distribution, hence this identity applies in the general case. I think this is the nicest result of the post, even though it's quite simple.
- ^
A latent with mediation error
has total remainder at least , since is decreasing. So, the floor applies at the budget, not just on it. - ^
I suspect this is part of why minds and systems function with discrete (digital) schemas to represent information: symbolic language, error correction, the genetic code, etc.
- ^
The slope of the curve has a closed form:
, so the marginal exchange rate at any point is , where is the correlation not yet explained. The going rate is always the corner rate of the residual. At the right end the residual is all of , giving ; at the left end the residual goes to zero and the rate goes to zero with it, which is the infinite slope at . The agreement a marginal shared latent explains scales with the residual correlation, while the remainder it incurs scales with how different the views still are. - ^
is the limit of the diagonal crossing, i.e. the worst case over sources; at any fixed the balanced optimum is somewhat cheaper. - ^
The runs are exchangeable rather than Gaussian, so this is reassurance that the half-bit is not an artifact of our Gaussian setting. It turns out to be a Fisher information statement (doubling data halves variance) that works asymptotically for any regular parametric family.
- ^
The curve exists for every distribution, not just Gaussians. The sum rule is distribution-free, so the minimal total remainder at mediation budget
is , where is the relaxed Wyner common information of the pair. This is defined for any two views, but there is no closed form in the general case. Its left endpoint is already informative in general: at exact mediation the total remainder is , which is positive for generic pairs. This is the general-case analog of the half-bit. - ^
This is speculative: I'm not too sure how useful it is to actually compute remainders for practical purposes. I hope to write more on this later.
- ^
The remaining hard case for the general bounty lives near decomposable distributions (tiny-mixtures territory), which Gaussian methods provably can't reach. Again, maybe more on this later.
相关文章
- 鹅鸭杀手游布谷鸟怎么玩-布谷鸟玩法教学 08-04
- 《明日方舟:终末地》六大毕业配队介绍 08-04
- 燕云十六声最终BOSS是谁 燕云妙善大师田英打法 08-04
- 鹅鸭杀加拿大鹅是干嘛的-加拿大鹅技能效果介绍 08-04
- 小触控连点器如何使用 08-04
- 无限暖暖不思议亡骨之宴任务怎么做-不思议亡骨之宴任务流程攻略 08-04