The short version
I spent three weeks on one idea: rip the learned position table out of CLIP-style models and use rotary position embeddings (RoPE) instead. Then a literature review I designed to kill my own paper killed three of my claims.
- 2D RoPE gives a SigLIP ViT-B +2.5 ImageNet zero-shot at fixed data, fixed compute and identical parameter count, with retrieval R@5 up 4 to 5 points, and the ImageNet gain grows to +3.0 at ViT-L. Three seeds on anything I claim.
- A RoPE text tower trained at 77 tokens reads captions about twice that long with retrieval staying nearly flat, and the only thing that changes at test time is one line of arithmetic on the rotation frequencies.
- TULIP (ICLR 2025) had already built a RoPE text tower for CLIP with the same frequency trick and beaten Long-CLIP. Two 2025 VLM tech reports had already put 2D RoPE inside a SigLIP encoder. And my theory section restated a 2021 theorem.
- What survived is the measurement. Nobody had run the clean ablation, and one of those tech reports publishes numbers pointing the other way.
How a transformer knows where things are
Attention has no sense of order: the dot product between a query and a key has no idea whether the two tokens are neighbors or fifty apart. CLIP patches this on both towers with a learned absolute position table, one trainable vector per position. The model learns what tends to sit at position 37 without ever learning how 37 relates to 36, and the table has a fixed 77 rows on the text side, so past the cap it can't run at all. The cap is the shape of a parameter, so you can't turn it up.
What RoPE does instead
RoPE came from RoFormer in 2021. Instead of adding a position vector to the token, it rotates the query and the key by an angle proportional to their position, right before the dot product. Pair a head's channels into d/2 planes, rotate a token at position m by m·theta in each one, and because rotations add their angles the score depends only on the offset n-m: two tokens three apart look the same at (0, 3) and at (500, 503). Length is preserved, so only phase carries position.
One detail matters later. Each plane spins at its own rate, theta_j = b-2j/d with base b = 10,000: a clock with d/2 hands, where the fast ones turn about a radian per token and resolve local order and the slow ones take thousands of tokens per turn.
The same operation, on a grid
A text token has one index, an image patch has two, so the standard trick, used by RoPE-ViT and EVA, is axial: rotate half of each head's channels by the row index and half by the column, and the score sees the two offsets separately, so 2D RoPE ends up being one 1D RoPE per axis.
In my OpenCLIP fork the Rotary2D module is about 15 lines applied to q and k in every attention layer, and it adds zero parameters, so every number below is measured at exactly equal parameter count. Feed it torch.arange(seq_len), skip the channel split, and the same module is 1D text RoPE. Two things I could test:
- P1: 2D RoPE should improve a sigmoid-loss (SigLIP) image tower at fixed budget, and the gain should be bigger on retrieval, which needs fine-grained spatial alignment, than on classification.
- P2: a 1D RoPE text tower trained on short captions should still read longer ones at inference, because position gets computed instead of looked up.
The setup
A word on the loss, because my novelty claim hung on it. CLIP trains with InfoNCE, a softmax over every caption in the batch, so the model picks its caption out of a lineup. SigLIP (ICCV 2023) swaps that for an independent yes/no decision per pair, with no normalization across the batch. RoPE was already inside EVA-CLIP's InfoNCE tower, present but never ablated, and I couldn't find published evidence on whether a positional bias tuned under one loss carries over to the other.
I forked OpenCLIP and made one methodological bet early. Every run is a rung on an additive ladder, and each rung differs from the one below it by exactly one ingredient: InfoNCE, SigLIP, +qk-norm, +2D RoPE, +attention pooling, +focal sigmoid. Where a stock config would have changed two things at once, I split it in two. It's slower than running whatever looks interesting, and it's the only reason my causal claims survived what happened later.
Training: ViT-B/16, CC12M (about 10M pairs), roughly 3 epochs, fixed global batch, H100s. The RoPE rung showed up right away. ImageNet-1k top-1 and COCO text-to-image R@5, mean of three seeds:
| arm | IN-1k top-1 | COCO t→i R@5 |
|---|---|---|
| CLIP (InfoNCE) | 23.87 | 36.19 |
| SigLIP baseline | 24.13 | 36.35 |
| + 2D RoPE | 26.59 | 40.06 |
| + 2D RoPE + qk-norm | 26.70 | 40.31 |
The ladder had stacked RoPE on top of qk-norm, so I ran isolation arms with RoPE alone on plain SigLIP, three seeds, to check I wasn't looking at an interaction. I wasn't: adding qk-norm on top of RoPE moves ImageNet by 0.11, and per-seed noise goes up to 0.26. The focal-weighted sigmoid variant made things worse, and I kept it in the paper as a reported negative. Retrieval gained more than classification, which is what P1 predicted, and ViT-L/14 widened the gap: 25.42 to 28.45, +3.0, ahead on every metric on every seed.
The text side, and one experiment I threw away
P2 asks whether a model trained only on 77-token captions can read longer ones at inference. In the absolute tower, position is a lookup into a table with exactly 77 rows, so it can't even attempt it. The usual workaround, which Long-CLIP builds on, is to interpolate the table, stretching 77 learned vectors across 164 slots so each position gets a blend of its trained neighbors, which gets the positions back but puts every token at coordinates the model never saw.
RoPE computes position instead, so it runs at any length, and it has its own failure mode. The phase in plane j at position m is m·theta_j, so a model trained at 77 has never seen phases past 77·theta_j, and at 150 the slow planes hand it angles that never occurred in training. The fix, known in LLM work as NTK-aware scaling, is one arithmetic change at inference:
# NTK-aware rescaling: the whole "method". Inference only, no retraining. s = L_test / L_train # e.g. 164 / 77 = 2.13 eff_base = base * s ** (head_dim / (head_dim - 2)) # 10_000 -> ~21_800 at s = 2.13 inv_freq = eff_base ** (-torch.arange(0, head_dim, 2) / head_dim)
The exponent d/(d-2) makes the largest phase at the test length match the largest phase seen in training, so every rotation the model meets at 164 is one it already met at 77, and the fast planes lose a little resolution. The community found this for LLaMA in mid 2023 and YaRN wrote it up later, and I hadn't seen it tried on a contrastive text encoder.
My first design was broken in a useful way. I trained at 248 tokens and "extrapolated" to 320, then noticed the eval captions top out around 222 tokens, so past 248 the test set is 100% padding and a flat curve proves nothing. I binned it and rebuilt: both arms trained at 77 tokens, CLIP's real limit, evaluated at 77, 128, 164, 196 and 248, where 164 is about the p99 of actual caption content on these benchmarks.
Both arms are the standard CLIP text tower, identical down to the parameter (about 149.7M) except for the position code. Trained on DenseFusion-1M, evaluated on Urban-1k and DOCCI, with the metric being the change in Recall@1 against each arm's own 77-token score. Urban-1k text-to-image, mean of three seeds:
| text tower | ΔR@1 at L=164 | ΔR@1 at L=248 |
|---|---|---|
| Absolute + interpolation | −4.8 | −7.2 |
| RoPE, NTK off | −4.5 | −4.5 |
| RoPE, NTK on | −1.3 | −0.4 |
Two controls make me trust that table. The NTK-off row shows raw RoPE extrapolating about as badly as interpolation, so the base change is the ingredient that matters. And if I replace every token past position 77 with padding, RoPE+NTK drops from 17.9 to 7.7 R@1, so the model is really reading the extra text.
The paper, and one proposition too many
I wrote it up as A Little RoPE Goes a Long Way, with three claimed contributions: the two experiments above, plus a theory section I was pleased with. Every term in the sigmoid loss depends only on inner products between normalized embeddings, and those survive rotating the whole space by an orthogonal transform, so the loss pins down relative geometry and says nothing about the absolute frame. RoPE also injects position through rotations that preserve norms and encode only relative offsets, which I argued is why rotation is the right positional bias for contrastive training.
The kill step
Before submitting anywhere I ran an adversarial novelty review, briefed to kill the claims rather than confirm them: three parallel sweeps over RoPE in vision, long-context CLIP text and the theory, each required to fetch primary sources and name the closest prior work.
The theory died in minutes. Zimmermann et al. (ICML 2021) proved that InfoNCE-family minimizers recover the true latents up to an orthogonal transform, so the gauge freedom I thought I'd spotted is the central result of a four-year-old spotlight paper. The 1D/2D unification exists at least four times over, in RoPE-ViT, LieRE, STRING and ComRoPE. And my proposition is about one global rotation of the final embeddings, while RoPE rotates queries and keys inside attention, so the first says nothing about the second. I'd written an analogy and formatted it as a proposition.
P1's framing had been false for about six weeks. Abstract-level searches for "RoPE SigLIP" return nothing, which is why I believed my own abstract. The pairing lives inside VLM tech reports, where that search can't reach. ByteDance's Seed1.5-VL (May 2025) trains a 2D-RoPE ViT with the SigLIP loss. Kwai's Keye-VL (July 2025) adds 2D RoPE to a SigLIP encoder, continues pretraining with the SigLIP loss, and publishes a table where RoPE does not beat the non-RoPE baseline at standard resolution. Meta's Perception Encoder had individually ablated 2D RoPE under softmax CLIP in April 2025.
Then the second sweep returned a paper I'd never seen: TULIP: Token-length Upgraded CLIP, Najdenkoska et al., arXiv October 2024, accepted at ICLR 2025. Reading its method section was uncomfortable. It replaces CLIP's absolute text position embeddings with relative ones, "we implement P_g(i) as Rotary Positional encodings (RoPE)", uses NTK-aware scaling for the same reason I did, and evaluates at 248 tokens against Long-CLIP and wins. That's my P2, published before I ran my first experiment. The one real difference is that TULIP's length expansion is a trained pipeline, distillation plus a fine-tuning epoch on long captions, while mine happens at inference only. And the training-free half exists too: LongEmbed (EMNLP 2024) did training-free NTK on text-embedding retrieval models and argued that future embedding models should use RoPE.
My related-work section cited seven papers. The review surfaced five directly adjacent works I had not cited at all, one of them peer-reviewed and published before the project's first commit.
The scoreboard
- C1, rotation unification plus the invariance proof. Died. Both halves are published (Zimmermann 2021, and RoPE-ViT / LieRE / STRING / ComRoPE), and the argument connecting them doesn't hold as stated. Demoted to background.
- C2, RoPE under the sigmoid loss. Narrowly alive. Seed1.5-VL and Keye-VL got there first, as confounded tech-report evidence. Nobody has the controlled, from-scratch, fixed-budget, seed-replicated isolation, and Keye-VL's numbers point the other way. That contradiction is what my experiment resolves.
- C3, training-free long context via RoPE and NTK. Narrowly alive. TULIP owns RoPE text tower plus NTK plus beating Long-CLIP, LongEmbed owns training-free NTK for contrastive text, and what's left is the intersection: zero-retraining extension inside the multimodal tower, plus the NTK-off ablation and the corruption control that neither paper has.
What survived, and what I would do differently
The experiments survived and the claims didn't. Every number above is real and replicated, and as far as I can tell it's still the only controlled measurement of its kind. What died is the story I wrapped around the numbers, the "first to combine RoPE with X" framing, and what's left is the first clean isolation of RoPE under the sigmoid loss plus evidence that in a CLIP text tower a one-line rescale at inference does what TULIP needed a training pipeline for.
- Run the kill review before writing the abstract. It cost me a day and rewrote three contribution claims, and three weeks earlier it would have changed the experiments, because I'd have trained a TULIP-style fine-tuned arm as a baseline.
- Search where the field actually publishes. My hinge claim was falsified by two tech reports that abstract-level search can't see, so if your claim is "X has never been paired with Y", you have to grep the 80-page tech reports.
- "First to combine" is easy to lose, because combinations get published constantly, in passing, deep inside VLM reports. "First controlled measurement" is much harder to lose, because nobody scoops a clean ablation by accident.
- The experiment I threw in the bin was the most valuable one. Killing my own padding-contaminated eval is why the second version held up under review, and the habit that caught it is the same one that caught the novelty problem, just pointed outward too late.
The paper is being reframed, with an honest venue target. The full adversarial review, with per-claim verdicts and every citation, is in NOVELTY_REVIEW.md in the repo.
References
- Su et al., RoFormer: Enhanced Transformer with Rotary Position Embedding, 2021. arXiv:2104.09864
- Heo et al., Rotary Position Embedding for Vision Transformer, ECCV 2024. arXiv:2403.13298
- Zhai et al., Sigmoid Loss for Language Image Pre-Training (SigLIP), ICCV 2023. arXiv:2303.15343
- Najdenkoska et al., TULIP: Token-length Upgraded CLIP, ICLR 2025. arXiv:2410.10034
- Zhu et al., LongEmbed: Extending Embedding Models for Long Context Retrieval, EMNLP 2024. arXiv:2404.12096
- Zhang et al., Long-CLIP: Unlocking the Long-Text Capability of CLIP, ECCV 2024. arXiv:2403.15378
- Zimmermann et al., Contrastive Learning Inverts the Data Generating Process, ICML 2021. arXiv:2102.08850
- Bolya et al., Perception Encoder, 2025. arXiv:2504.13181
- Seed1.5-VL Technical Report, 2025. arXiv:2505.07062
- Kwai Keye-VL Technical Report, 2025. arXiv:2507.01949
- Peng et al., YaRN: Efficient Context Window Extension of Large Language Models, 2023. arXiv:2309.00071