Add Autoconfig, Coordinated_Optimizer and Sharding keras implementations for Tensor Parallel Autosharding #21707

buildwithsuhana · 2025-10-02T10:19:24Z

This PR introduces support for tensor parallelism autosharding in Keras, enabling users to shard large model layers across multiple devices. This is a crucial feature for training models that are too large to fit into the memory of a single accelerator.

The implementation is centered around two new components:

autoconfig.py: This module contains the logic to analyze a Keras model, identify sharding candidates (e.g., Dense, EinsumDense layers), and generate a sharding plan.

coordinated_optimizer.py: This is an optimizer wrapper that consumes the sharding plan. During training, it intercepts gradients for sharded variables and performs a collective AllReduce to ensure weight updates are correctly synchronized across all devices.

…hsuhana/keras into Tensor_parallel_keras_2

gemini-code-assist · 2025-10-02T10:19:41Z

Summary of Changes

Hello @buildwithsuhana, I'm Gemini Code Assist¹! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly enhances Keras's capabilities for large-scale model training by introducing foundational support for tensor parallelism autosharding. It provides mechanisms to automatically determine how model layers should be split across multiple devices and a specialized optimizer to manage the distributed training process, including sharding optimizer states and synchronizing gradients. This enables users to train models that exceed the memory capacity of a single accelerator, making distributed training more accessible and efficient within the Keras ecosystem.

Highlights

Automatic Sharding Configuration: Introduced "autoconfig.py" to intelligently analyze Keras models (e.g., Dense, EinsumDense, Embedding layers) and generate a tensor parallelism sharding plan, classifying Dense layers as "up-projection" or "down-projection" for optimal splitting.
Coordinated Optimizer for Distributed Training: Added "coordinated_optimizer.py" which provides "CoordinatedOptimizer" for managing sharded optimizer states and synchronizing gradients across devices, and "TensorParallelOptimizer" as a Keras-compatible wrapper.
Gradient Synchronization Logic: The "CoordinatedOptimizer" includes logic to perform "all-reduce" operations on gradients of column-parallel sharded weights, ensuring correct updates in a distributed setting.
Optimizer State Sharding: Implemented functionality to partition optimizer state variables (like momentum and velocity) across multiple devices, reducing memory footprint per device.
Comprehensive Testing: New test files ("autoconfig_test.py", "coordinated_optimizer_test.py") were added to validate the automatic sharding configuration and the distributed optimizer's behavior, including handling of replicated and sharded states, and serialization.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature	Command	Description
Code Review	`/gemini review`	Performs a code review for the current pull request in its current state.
Pull Request Summary	`/gemini summary`	Provides a summary of the current pull request in its current state.
Comment	@gemini-code-assist	Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help	`/gemini help`	Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

gemini-code-assist

Code Review

This pull request introduces significant new functionality for tensor parallelism autosharding in Keras, including modules for automatic configuration and a coordinated optimizer. The implementation is well-structured, with new logic for analyzing models, generating sharding plans, and synchronizing gradients. However, I've identified a few issues that need attention. There is a critical bug in the CoordinatedOptimizer where a method for applying gradients with sharded states is called but not defined. I also found a couple of high-severity issues related to incorrect logic for matching optimizer states and gathering sharded parameters, which could lead to runtime errors or incorrect behavior. Additionally, there are some medium-severity issues regarding code clarity, such as unused parameters. The accompanying tests are a good start but do not cover the code path with the critical bug.

keras/src/distribution/tensor_parallel/coordinated_optimizer.py

keras/src/distribution/tensor_parallel/sharding_keras.py

keras/src/distribution/tensor_parallel/autoconfig.py

keras/src/distribution/tensor_parallel/coordinated_optimizer.py

buildwithsuhana added 6 commits October 1, 2025 15:59

adding autoconfig and coordinated_optimizer

dd3181e

Reformatting

bcae2f6

Added sharding keras

439643b

Merge branch 'keras-team:master' into Tensor_parallel_keras_2

36edcb9

Reformatting files

b7862d9

Merge branch 'Tensor_parallel_keras_2' of https://github.com/buildwit…

e8b51f7

…hsuhana/keras into Tensor_parallel_keras_2

google-ml-butler bot added the size:XL label Oct 2, 2025

google-ml-butler bot assigned gbaned Oct 2, 2025

gemini-code-assist bot reviewed Oct 2, 2025

View reviewed changes

buildwithsuhana marked this pull request as draft October 2, 2025 17:28

buildwithsuhana added 5 commits October 3, 2025 11:55

Reformatting according to changes in distributed_backend

3383dec

Reformatting according to changes in distributed_backend

5824c66

Refactoring the code

9cf5c7f

refactoring

996a154

refactoring

31994da

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Add Autoconfig, Coordinated_Optimizer and Sharding keras implementations for Tensor Parallel Autosharding #21707

Add Autoconfig, Coordinated_Optimizer and Sharding keras implementations for Tensor Parallel Autosharding #21707

Uh oh!

buildwithsuhana commented Oct 2, 2025

Uh oh!

gemini-code-assist bot commented Oct 2, 2025

Uh oh!

gemini-code-assist bot left a comment

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Add Autoconfig, Coordinated_Optimizer and Sharding keras implementations for Tensor Parallel Autosharding #21707

Are you sure you want to change the base?

Add Autoconfig, Coordinated_Optimizer and Sharding keras implementations for Tensor Parallel Autosharding #21707

Uh oh!

Conversation

buildwithsuhana commented Oct 2, 2025

Uh oh!

gemini-code-assist bot commented Oct 2, 2025

Summary of Changes

Highlights

Footnotes

Uh oh!

gemini-code-assist bot left a comment

Choose a reason for hiding this comment

Code Review

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!