Meta has open-sourced Rebalancer, a C++ library with a Python interface for solving assignment problems. It decides which objects go into which bins under constraints and objectives. According to the Engineering at Meta’s post, Rebalancer has handled resource allocation across Meta for over 9 years. The release ships under Apache 2.0 with documentation, a PyPI package and a debugging UI called Rebalancer Explorer.
Is it deployable? Yes. pip install rebalancer installs v1.0.4 for Python 3.12+, with prebuilt wheels for Linux x86-64 and macOS 14+ ARM64. .deb, .rpm and Homebrew packages also exist. PyPI still classifies the project as Alpha.
What Problem Does Rebalancer Solve?
Assignment problems show up across Meta’s stack. Racks go into datacenters, servers go to services, tasks go to servers, and user traffic goes to datacenters. Meta names 2 blockers: usability and scalability. Engineers struggle to turn policies into precise formulas, and many problems are NP-hard and too large for commercial solvers.
Rebalancer’s answer is to separate how a problem is specified from how it is solved. The design is detailed in the OSDI 2024 paper, Optimizing Resource Allocation in Hyperscale Datacenters.
How the Specification Layer Works
The spec language has 3 layers:
- Modeling constructs: dimensions (attributes such as CPU or storage), partitions (groups of objects), scopes (groups of bins) and utilization.
- Expression API: aggregate utilization with SUM or MAX, or transform it with operations such as SQUARE.
- Spec API: dozens of predefined objectives and constraints, listed in the docs.
Meta’s example models tasks as objects, servers as bins and racks as a scope. A CapacitySpec caps CPU and storage per server. A GroupCountSpec keeps 1 job type per rack. A BalanceSpec balances each server’s utilization across both dimensions.
One Expression Graph, Two Solvers
Rebalancer compiles the spec into a directed acyclic expression graph. Leaf nodes hold utilization values; aggregation and transformation nodes sit above them. Users supply an initial assignment and a stopping condition. Constraints that the initial assignment already violates become high-priority goals.
Optimal solver: The graph is translated into a mixed integer program for FICO Xpress, Gurobi or HiGHS. Variable aggregation and symmetry breaking shrink models. The worst-case model size is still O(objects × bins). Meta’s largest problems are too big for any MIP solver.
Local search: This solver works directly on the expression graph. It explores moves of objects to other bins, with a worst-case neighborhood of O(objects + bins). It then applies the best candidate that breaks no constraint. Evaluation is parallelized, reaching millions of evaluations per second, and the search space is pruned.
Meta uses local search for almost all large problems and MIP for small to mid-size ones, often prototyping with MIP first.
- About 40 million assignment problems solved per day, across 30+ unique formulations.
- P99 solve time of 12 seconds on 265k objects and 3.2k bins.
- Problems above 1 million objects and 5k bins average 171 seconds, across 3.4k+ runs.
Best Use Cases for Rebalancer
- Placing shards, tasks or containers on a cluster: Assign work to servers under CPU and memory caps while spreading replicas across racks. Meta’s Shard Manager and RAS run this pattern.
- Balancing traffic and workloads across regions: Route user traffic or jobs to datacenters, trading latency against load. Taiji does this for edge traffic, and Meta balances ML training by priority.
- Operational assignment outside infrastructure: Map support tickets to engineers, meetings to rooms or desks to people under capacity rules. Meta has done all 3.
Debugging With Rebalancer Explorer
Modelers at Meta spent most of their time debugging solver behavior. Rebalancer Explorer is a Dockerized web UI built for this. It shows binding constraints, relaxation effects, and why an object landed in a bin.

