By using this site, you agree to the Privacy Policy and Terms of Use.
Accept
AIModelKitAIModelKitAIModelKit
  • Home
  • News
    NewsShow More
    SpaceXAI’s Grok Tool Uploading Users’ Entire Codebase to Cloud Storage: What You Need to Know
    SpaceXAI’s Grok Tool Uploading Users’ Entire Codebase to Cloud Storage: What You Need to Know
    4 Min Read
    New York Leads the Way: First State to Enforce One-Year Moratorium on New AI Data Centers
    New York Leads the Way: First State to Enforce One-Year Moratorium on New AI Data Centers
    4 Min Read
    AI Replacing New York Nurses: Why Patients Should be Concerned About Quality of Care
    AI Replacing New York Nurses: Why Patients Should be Concerned About Quality of Care
    5 Min Read
    Navigating AI Agent Crawlers and Cloudflare’s New Rules: A Comprehensive Guide
    Navigating AI Agent Crawlers and Cloudflare’s New Rules: A Comprehensive Guide
    5 Min Read
    How Apple’s Self-Driving Car Program Paved the Way for Advanced AI Chip Technology
    How Apple’s Self-Driving Car Program Paved the Way for Advanced AI Chip Technology
    4 Min Read
  • Open-Source Models
    Open-Source ModelsShow More
    GlucoFM: Advanced Foundation Model for Continuous Glucose Monitoring Insights
    GlucoFM: Advanced Foundation Model for Continuous Glucose Monitoring Insights
    5 Min Read
    AgentHands: Creating Interactive Hand Gestures for Enhanced Conversations with Spatially Grounded Agents in XR
    AgentHands: Creating Interactive Hand Gestures for Enhanced Conversations with Spatially Grounded Agents in XR
    5 Min Read
    Exploring How Mobility Enhances Language Models’ Understanding of Location
    Exploring How Mobility Enhances Language Models’ Understanding of Location
    5 Min Read
    Optimize Candidate Biomarkers with Our AI Tool for Wearable Sensor Data Analysis
    Optimize Candidate Biomarkers with Our AI Tool for Wearable Sensor Data Analysis
    4 Min Read
    Beyond BMI: Assessing Cardiometabolic Risk Using Smartphone Images
    Beyond BMI: Assessing Cardiometabolic Risk Using Smartphone Images
    5 Min Read
  • Guides
    GuidesShow More
    Your Comprehensive Guide to Practical Constraint Decoding: Basics and Applications
    Your Comprehensive Guide to Practical Constraint Decoding: Basics and Applications
    6 Min Read
    KDnuggets Weekly Data Science News Roundup: Highlights from July 20, 2026
    KDnuggets Weekly Data Science News Roundup: Highlights from July 20, 2026
    4 Min Read
    Unlock Your AI Potential with Kaggle and Google’s Free 5-Day Agentic AI Course
    Unlock Your AI Potential with Kaggle and Google’s Free 5-Day Agentic AI Course
    6 Min Read
    Top 5 High-Performance MCP Servers for Optimal Agentic Development
    Top 5 High-Performance MCP Servers for Optimal Agentic Development
    6 Min Read
    Top 5 Free Resources for Understanding Agentic AI: Unlock Your Knowledge
    Top 5 Free Resources for Understanding Agentic AI: Unlock Your Knowledge
    6 Min Read
  • Tools
    ToolsShow More
    Unlock Agentic Coding: Experimenting with Qwen 3.8-Flash-Next on NVIDIA GB300 NVL72
    Unlock Agentic Coding: Experimenting with Qwen 3.8-Flash-Next on NVIDIA GB300 NVL72
    6 Min Read
    Unlock Agentic Coding: Experimenting with Qwen 3.8 Flash-Next 176B Model on NVIDIA GB300 NVL72
    Unlock Agentic Coding: Experimenting with Qwen 3.8 Flash-Next 176B Model on NVIDIA GB300 NVL72
    5 Min Read
    Optimizing LFM2.5 Q4_0 Checkpoints through Quantization-Aware Distillation Techniques
    Optimizing LFM2.5 Q4_0 Checkpoints through Quantization-Aware Distillation Techniques
    4 Min Read
    Deploy Qwen 3.8-2.4T-A95B: A Configurable 2.4T Parameter Model on NVIDIA GB300 NVL72 for Enhanced Reasoning
    Deploy Qwen 3.8-2.4T-A95B: A Configurable 2.4T Parameter Model on NVIDIA GB300 NVL72 for Enhanced Reasoning
    6 Min Read
    Optimize Your AI Models with Baseten on Hugging Face Inference Providers 🔥
    Optimize Your AI Models with Baseten on Hugging Face Inference Providers 🔥
    5 Min Read
  • Events
    EventsShow More
    Exploring the Future of EdTech: Highlights from the ‘Best of ISTE’ Virtual Playground
    Exploring the Future of EdTech: Highlights from the ‘Best of ISTE’ Virtual Playground
    4 Min Read
    Empowering Veteran Students: Effective Teaching Strategies in Technology and Learning
    Empowering Veteran Students: Effective Teaching Strategies in Technology and Learning
    4 Min Read
    NVIDIA Partners with NSF to Enhance AI Research and Education Through State and Regional AI Hubs Across the US
    NVIDIA Partners with NSF to Enhance AI Research and Education Through State and Regional AI Hubs Across the US
    5 Min Read
    South Korea Unveils AI Future at AI Summit with NVIDIA and Strategic Partners
    South Korea Unveils AI Future at AI Summit with NVIDIA and Strategic Partners
    5 Min Read
    NVIDIA Launches First Open-Source GPU-Accelerated Framework for Medical Physics Simulations
    NVIDIA Launches First Open-Source GPU-Accelerated Framework for Medical Physics Simulations
    5 Min Read
  • Ethics
    EthicsShow More
    AI Giants Warn: Impending Cybersecurity Crisis Looms in Just Months
    AI Giants Warn: Impending Cybersecurity Crisis Looms in Just Months
    6 Min Read
    Assessing the Environmental Impact of Data Centres: Are We Finally Acknowledging the Consequences?
    Assessing the Environmental Impact of Data Centres: Are We Finally Acknowledging the Consequences?
    5 Min Read
    Survey Reveals Surprising Impact of AI on Job Losses: Insights from Workers
    Survey Reveals Surprising Impact of AI on Job Losses: Insights from Workers
    6 Min Read
    Understanding DAO-to-DAO Voting: On-Chain and Off-Chain Mechanisms Explored
    Understanding DAO-to-DAO Voting: On-Chain and Off-Chain Mechanisms Explored
    5 Min Read
    Taiwan Prosecutes Nine Individuals for Smuggling Advanced AI Servers to China: A Tech Industry Update
    Taiwan Prosecutes Nine Individuals for Smuggling Advanced AI Servers to China: A Tech Industry Update
    4 Min Read
  • Comparisons
    ComparisonsShow More
    InternBootcamp: Enhancing LLM Reasoning Through Verifiable Task Scaling Techniques
    InternBootcamp: Enhancing LLM Reasoning Through Verifiable Task Scaling Techniques
    4 Min Read
    Enhancing Anomaly Detection in Collider Experiments through Contrastive Learning for Better Interpretability
    Enhancing Anomaly Detection in Collider Experiments through Contrastive Learning for Better Interpretability
    6 Min Read
    Exploring the Impact of Quantization on Self-Explanations in Large Language Models: Can LLMs Explain Themselves?
    Exploring the Impact of Quantization on Self-Explanations in Large Language Models: Can LLMs Explain Themselves?
    5 Min Read
    CytoNet: A Foundation Model for Understanding the Human Cerebral Cortex at Cellular Resolution
    CytoNet: A Foundation Model for Understanding the Human Cerebral Cortex at Cellular Resolution
    5 Min Read
    Optimizing Nonconvex-Nonconcave Min-Max Problems with a Limited Maximization Domain: Insights from [2110.03950]
    Optimizing Nonconvex-Nonconcave Min-Max Problems with a Limited Maximization Domain: Insights from [2110.03950]
    5 Min Read
Search
  • Privacy Policy
  • Terms of Service
  • Contact Us
  • FAQ / Help Center
  • Advertise With Us
  • Latest News
  • Model Comparisons
  • Tutorials & Guides
  • Open-Source Tools
  • Community Events
© 2025 AI Model Kit. All Rights Reserved.
Reading: Why SAEs Trained on Identical Data Sets Can Discover Different Features
Share
Notification Show More
Font ResizerAa
AIModelKitAIModelKit
Font ResizerAa
  • 🏠
  • 🚀
  • 📰
  • 💡
  • 📚
  • ⭐
Search
  • Home
  • News
  • Models
  • Guides
  • Tools
  • Ethics
  • Events
  • Comparisons
Follow US
  • Latest News
  • Model Comparisons
  • Tutorials & Guides
  • Open-Source Tools
  • Community Events
© 2025 AI Model Kit. All Rights Reserved.
AIModelKit > Comparisons > Why SAEs Trained on Identical Data Sets Can Discover Different Features
Comparisons

Why SAEs Trained on Identical Data Sets Can Discover Different Features

aimodelkit
Last updated: April 13, 2025 8:45 am
aimodelkit
Share
Why SAEs Trained on Identical Data Sets Can Discover Different Features
SHARE

Understanding the Impact of Random Initializations on TopK Sparse Autoencoders

In the realm of machine learning, particularly in the training of neural networks, the role of initialization cannot be overstated. This article delves into the intriguing findings from our investigation of TopK Sparse Autoencoders (SAEs), specifically how variations in random initialization can lead to divergent feature representations even when trained on identical datasets with the same batch order.

Contents
  • Divergence in Latent Representations
  • Interpretability of Unshared Latents
  • Feature Splitting and Absorption
  • Stability Across Different Architectures
  • Methodology: Measuring Latent Alignment
  • Latent Overlap Across Multiple Models
  • Frequency of Latent Activation
  • The Influence of SAE Size on Feature Overlap
  • Investigating Interpretability of Unique Latents
  • Conclusion

Divergence in Latent Representations

When two TopK SAEs are trained using the same data but with different random initializations, a fascinating phenomenon occurs. Our study reveals that only about 53% of the features are shared between these two models. This relatively low overlap suggests that a significant number of latents in one SAE do not have a close counterpart in the other, and vice versa. The implication here is profound: the features learned by SAEs are not fixed or universally applicable, but rather can be highly variable based on initialization.

Interpretability of Unshared Latents

Interestingly, many of the unshared latents exhibit interpretability. This raises the question of how different training paths can lead to distinct yet interpretable representations. Furthermore, we observed that narrower SAEs tend to have a higher overlap of features across random seeds. In contrast, as the size of the SAE increases, the degree of overlap diminishes. This trend aligns with existing literature on feature splitting and absorption, indicating that the features learned by SAEs can be somewhat arbitrary.

Feature Splitting and Absorption

The behavior of SAEs supports the idea that learned features are not atomic. Instead, different configurations can lead to various interpretations of the same latent features. As the size of the SAEs increases, we also see a phenomenon known as feature absorption, where some latents gain an “implicit” meaning alongside their “explicit” feature interpretation. This duality in representation can allow models to learn disjoint representations even when they are trained on the same data.

Stability Across Different Architectures

Our findings suggest that the architecture of the SAE plays a crucial role in the stability of feature learning under different random seeds. Previous studies have indicated that certain architectures, like ReLU SAEs trained with an L1 penalty, show significant stability across different initializations. In contrast, TopK SAEs appear to benefit from methods that align different seeds, highlighting the need for careful consideration in architectural choices.

More Read

OpenAI Unveils GPT-4.1 Family: Improved Performance and Long-Context Capabilities
OpenAI Unveils GPT-4.1 Family: Improved Performance and Long-Context Capabilities
Open-World Evaluation Techniques for Diverse Perspective Retrieval: Insights from Research 2409.18110
Scalable LLM Accelerator Fault Assessment: A Reinforcement Learning Approach
Enhancing Out-of-Distribution Detection in Autonomous Vessels Using Digital Twin Technology
Cloudflare Discovers Query Planning Bottleneck in ClickHouse Performance

Methodology: Measuring Latent Alignment

To quantify the alignment between independently trained SAEs, we employed the Hungarian algorithm. This method efficiently computes the matching between latents, maximizing the average cosine similarity between matched encoder and decoder vectors. The resulting alignment score provides a clear measure of how similarly the two models interpret the latent space.

Upon analyzing the distribution of cosine similarities, we observed that there are two distinct modes: one reflecting high similarity and another indicating low similarity. This duality suggests that while some latents are closely aligned, others diverge significantly. In cases where the encoder and decoder matchings disagree, the cosine similarity tends to be lower, reinforcing the complexity of the latent space.

Latent Overlap Across Multiple Models

Further exploration revealed that when introducing a third SAE trained with a different random seed, the overlap of shared latents decreased from 47% to 35%. This finding indicates that the majority of shared latents between the first two models also persist in their relationship with the third model, showcasing an interesting dynamic of latent retention across different configurations.

Frequency of Latent Activation

An important aspect of our investigation was examining the frequency of latent activation across models. We found that the latents most frequently activated in SAE 1 were also those shared with SAE 2 and SAE 3. Conversely, the latents that activated infrequently in SAE 1 were those unique to that model. Intriguingly, some latents exclusive to SAE 1 exhibited a higher average firing rate than those present across all models, hinting at a complex relationship between activation frequency and latent representation.

The Influence of SAE Size on Feature Overlap

Our research also underscores a clear relationship between the size of the SAE and the fraction of unshared latents. Even when applying a more lenient metric for defining shared features, it became evident that larger SAEs retained a greater number of unique features. The computational demands of analyzing these larger models are significant, with implementations taking considerable time and resources, further complicating the exploration of this relationship.

Investigating Interpretability of Unique Latents

To delve deeper into the interpretability of unshared latents, we utilized an auto-interp approach to evaluate over 7,000 latents from two 32,768 latent SAEs. Our findings indicated a promising average interpretability score of 0.72, with a significant number of explanations falling within a reasonable range of clarity. However, the latents with low interpretability scores often correlated with low similarity across different seeds, suggesting that while some latents may be unique to a specific initialization, they might not lend themselves to clear interpretation.

Conclusion

The exploration of TopK SAEs trained under varying random initializations reveals a rich tapestry of latent representations that diverge significantly based on initial conditions. Our results challenge the notion of a universal set of features, highlighting the importance of viewing feature discovery as a compositional problem. As we continue to investigate these phenomena, we anticipate further insights into the intricate dance between architecture, initialization, and feature representation in neural networks.

By understanding these dynamics, we can better harness the capabilities of SAEs and other machine learning models, paving the way for more robust and interpretable AI systems.

Scalable Solutions Driven by Expert Domain Knowledge
Uncovering Position Bias and Ceiling Effects: A Permutation Diagnostic for Evaluating LLM Benchmarks
Anthropic Uncovers Three Key Infrastructure Bugs Affecting Claude’s Performance
AI-Assisted Development: Exploring Real-World Patterns, Common Pitfalls, and Ensuring Production Readiness – A Comprehensive Article Series
QConSF 2025: Accelerating Claude Code Development at Anthropic with AI Innovations

Sign Up For Daily Newsletter

Get AI news first! Join our newsletter for fresh updates on open-source models.

By signing up, you agree to our Terms of Use and acknowledge the data practices in our Privacy Policy. You may unsubscribe at any time.
Share This Article
Facebook Copy Link Print
Previous Article Advanced Bilingual French-English Language Model for Enhanced Communication Advanced Bilingual French-English Language Model for Enhanced Communication
Next Article Enhancing Explainable Moral Judgment Through Contrastive Ethical Insights from Large Language Models Enhancing Explainable Moral Judgment Through Contrastive Ethical Insights from Large Language Models

Stay Connected

XFollow
PinterestPin
TelegramFollow
LinkedInFollow

							banner							
							banner
Explore Top AI Tools Instantly
Discover, compare, and choose the best AI tools in one place. Easy search, real-time updates, and expert-picked solutions.
Browse AI Tools

Latest News

AI Giants Warn: Impending Cybersecurity Crisis Looms in Just Months
AI Giants Warn: Impending Cybersecurity Crisis Looms in Just Months
Ethics
Assessing the Environmental Impact of Data Centres: Are We Finally Acknowledging the Consequences?
Assessing the Environmental Impact of Data Centres: Are We Finally Acknowledging the Consequences?
Ethics
Exploring the Future of EdTech: Highlights from the ‘Best of ISTE’ Virtual Playground
Exploring the Future of EdTech: Highlights from the ‘Best of ISTE’ Virtual Playground
Events
Survey Reveals Surprising Impact of AI on Job Losses: Insights from Workers
Survey Reveals Surprising Impact of AI on Job Losses: Insights from Workers
Ethics
//

Leading global tech insights for 20M+ innovators

Quick Link

  • Latest News
  • Model Comparisons
  • Tutorials & Guides
  • Open-Source Tools
  • Community Events

Support

  • Privacy Policy
  • Terms of Service
  • Contact Us
  • FAQ / Help Center
  • Advertise With Us

Sign Up for Our Newsletter

Get AI news first! Join our newsletter for fresh updates on open-source models.

AIModelKitAIModelKit
Follow US
© 2025 AI Model Kit. All Rights Reserved.
Welcome Back!

Sign in to your account

Username or Email Address
Password

Lost your password?