The idea in plain language.
Agreement between embeddings could express which image parts belong to the same object.
How it works
The paper proposes representing a visual scene through vectors that agree at different levels of abstraction. Groups of similar vectors would indicate that image locations belong to a common part or object. This connects ideas from capsules, attention, contrastive learning, and neural fields. The contribution is a proposed representation scheme, rather than a completed working system.
What to keep in mind
The paper explicitly presents a conceptual system. Treat its design as a research hypothesis, not as demonstrated general visual understanding.
Source: How to represent part-whole hierarchies in a neural network. The original manuscript contains the methods, experiments, figures, and references. An arXiv posting date may follow an earlier conference publication. Read the linked record for version history.