Git Re-Basin: Merging Models modulo Permutation Symmetries
AI-generated Key Points
- Deep learning success attributed to solving complex non-convex optimization problems easily
- Simple algorithms like stochastic gradient descent effective in fitting large neural networks
- Neural network loss landscapes typically contain a single basin, considering permutation symmetries of hidden units
- Introduction of three algorithms to permute units and align them with reference model
- Transformation results in functionally equivalent weights in approximately convex basin near reference model
- Experimental demonstration of single basin phenomenon across various model architectures and datasets
- Discovery of zero-barrier linear mode connectivity between independently trained ResNet models on CIFAR-10 and CIFAR-100 datasets
- Independently trained networks can have meaningful differences in learned features under certain scenarios
- Model width and training time affect mode connectivity across different models and datasets
- Shortcomings of single basin theory discussed, including counterexample to linear mode connectivity hypothesis
- Study provides insights into optimization landscape of neural networks and understanding convergence towards similar solutions
Authors: Samuel K. Ainsworth, Jonathan Hayase, Siddhartha Srinivasa
Abstract: The success of deep learning is thanks to our ability to solve certain massive non-convex optimization problems with relative ease. Despite non-convex optimization being NP-hard, simple algorithms -- often variants of stochastic gradient descent -- exhibit surprising effectiveness in fitting large neural networks in practice. We argue that neural network loss landscapes contain (nearly) a single basin, after accounting for all possible permutation symmetries of hidden units. We introduce three algorithms to permute the units of one model to bring them into alignment with units of a reference model. This transformation produces a functionally equivalent set of weights that lie in an approximately convex basin near the reference model. Experimentally, we demonstrate the single basin phenomenon across a variety of model architectures and datasets, including the first (to our knowledge) demonstration of zero-barrier linear mode connectivity between independently trained ResNet models on CIFAR-10 and CIFAR-100. Additionally, we identify intriguing phenomena relating model width and training time to mode connectivity across a variety of models and datasets. Finally, we discuss shortcomings of a single basin theory, including a counterexample to the linear mode connectivity hypothesis.
Ask questions about this paper to our AI assistant
You can also chat with multiple papers at once here.
Assess the quality of the AI-generated content by voting
Score: 0
Why do we need votes?
Votes are used to determine whether we need to re-run our summarizing tools. If the count reaches -10, our tools can be restarted.
The previous summary was created more than a year ago and can be re-run (if necessary) by clicking on the Run button below.
Similar papers summarized with our AI tools
Navigate through even more similar papers through a
tree representationLook for similar papers (in beta version)
By clicking on the button above, our algorithm will scan all papers in our database to find the closest based on the contents of the full papers and not just on metadata. Please note that it only works for papers that we have generated summaries for and you can rerun it from time to time to get a more accurate result while our database grows.
Disclaimer: The AI-based summarization tool and virtual assistant provided on this website may not always provide accurate and complete summaries or responses. We encourage you to carefully review and evaluate the generated content to ensure its quality and relevance to your needs.