Transformers know more than they can tell -- Learning the Collatz sequence

arXiv — cs.LG•Monday, November 17, 2025 at 5:00:00 AM

NeutralArtificial Intelligence

The research explores how transformer models predict long steps in the Collatz sequence, revealing varying accuracies based on the encoding base. Models reached up to 99.7% accuracy for certain bases, indicating a strong capability in handling complex arithmetic functions.
This development is significant as it showcases the potential of transformer models in understanding intricate mathematical sequences, which could enhance their application in fields requiring advanced computational skills.
While no related articles were identified, the findings underscore the importance of model accuracy and learning patterns in AI, particularly in complex arithmetic tasks.

— via World Pulse Now AI Editorial System

Read Original

Was this article worth reading? Share it

Recommended Readings

arXiv — cs.CV2 days ago

Bridging Hidden States in Vision-Language Models

PositiveArtificial Intelligence

Vision-Language Models (VLMs) are emerging models that integrate visual content with natural language. Current methods typically fuse data either early in the encoding process or late through pooled embeddings. This paper introduces a lightweight fusion module utilizing cross-only, bidirectional attention layers to align hidden states from both modalities, enhancing understanding while keeping encoders non-causal. The proposed method aims to improve the performance of VLMs by leveraging the inherent structure of visual and textual data.

Read full article

via arXiv — cs.CV

arXiv — cs.LG2 days ago

Bias-Restrained Prefix Representation Finetuning for Mathematical Reasoning

PositiveArtificial Intelligence

The paper titled 'Bias-Restrained Prefix Representation Finetuning for Mathematical Reasoning' introduces a new method called Bias-REstrained Prefix Representation FineTuning (BREP ReFT). This approach aims to enhance the mathematical reasoning capabilities of models by addressing the limitations of existing Representation finetuning (ReFT) methods, which struggle with mathematical tasks. The study demonstrates that BREP ReFT outperforms both standard ReFT and weight-based Parameter-Efficient finetuning (PEFT) methods through extensive experiments.

Read full article

via arXiv — cs.LG

arXiv — cs.LG2 days ago

Multistability of Self-Attention Dynamics in Transformers

NeutralArtificial Intelligence

The paper titled 'Multistability of Self-Attention Dynamics in Transformers' explores a continuous-time multiagent model of self-attention mechanisms in transformers. It establishes a connection between self-attention dynamics and a multiagent version of the Oja flow, which computes the principal eigenvector of a matrix related to the value matrix in transformers. The study classifies the equilibria of the single-head self-attention system into four categories: consensus, bipartite consensus, clustering, and polygonal equilibria, noting that multiple stable equilibria can coexist.

Read full article

via arXiv — cs.LG

$$\pi$-Attention: Periodic Sparse Transformers for Efficient Long-Context Modeling$

arXiv — cs.CL2 days ago

$\pi$-Attention: Periodic Sparse Transformers for Efficient Long-Context Modeling

PositiveArtificial Intelligence

Read full article

via arXiv — cs.CL

arXiv — cs.LG2 days ago

Higher-order Neural Additive Models: An Interpretable Machine Learning Model with Feature Interactions

PositiveArtificial Intelligence

Higher-order Neural Additive Models (HONAMs) have been introduced as an advancement over Neural Additive Models (NAMs), which are known for their predictive performance and interpretability. HONAMs address the limitation of NAMs by effectively capturing feature interactions of arbitrary orders, enhancing predictive accuracy while maintaining interpretability, crucial for high-stakes applications. The source code for HONAM is publicly available on GitHub.

Read full article

via arXiv — cs.LG

arXiv — cs.CV2 days ago

RiverScope: High-Resolution River Masking Dataset

PositiveArtificial Intelligence

RiverScope is a newly developed high-resolution dataset aimed at improving the monitoring of rivers and surface water dynamics, which are crucial for understanding Earth's climate system. The dataset includes 1,145 high-resolution images covering 2,577 square kilometers, with expert-labeled river and surface water masks. This initiative addresses the challenges of monitoring narrow or sediment-rich rivers that are often inadequately represented in low-resolution satellite data.

Read full article

via arXiv — cs.CV