Lorenz Conditioned Networks¶
Lorenz Conditioned Networks (LCN) extends Pareto Conditioned Networks (PCN) by replacing the dominance criterion used to select and condition on past episodes. Where PCN uses Pareto dominance, LCN instead uses Lorenz dominance: each return’s objective values are sorted in increasing order and cumulatively summed (its Lorenz vector), and one return Lorenz-dominates another if its Lorenz vector is at least as high in every position. Lorenz dominance favors returns that distribute reward more equitably across objectives.
The lcn_lambda parameter controls λ-Lorenz dominance, an interpolation between strict Lorenz
dominance and ordinary dominance computed on the sorted returns:
fv = lcn_lambda * sorted(returns) + (1 - lcn_lambda) * lorenz_vector(returns)
lcn_lambda=0 recovers strict Lorenz dominance (maximizes fairness), while values closer to 1 relax
the fairness bias towards rewarding raw performance on the best-performing sorted objectives. Pick one of
the two distance_ref options to choose between the two:
"nondominated"(default): strict Lorenz dominance."lambda_lorenz": λ-Lorenz dominance, requireslcn_lambdato be set.
For more details, see:
Michailidis, D., Röpke, W., Roijers, D. M., Ghebreab, S., & Santos, F. P. (2026). Scalable Multi-Objective Reinforcement Learning with Fairness Guarantees using Lorenz Dominance. Journal of Artificial Intelligence Research, 85. https://jair.org/index.php/jair/article/view/19862
- class morl_baselines.multi_policy.lcn.lcn.LCN(env: Env | None, scaling_factor: ndarray, learning_rate: float = 0.01, gamma: float = 1.0, batch_size: int = 32, hidden_dim: int = 64, noise: float = 0.1, distance_ref: str = 'nondominated', lcn_lambda: float | None = None, project_name: str = 'MORL-Baselines', experiment_name: str = 'LCN', wandb_entity: str | None = None, log: bool = True, seed: int | None = None, device: device | str = 'auto', model_class: Type[BasePCNModel] | None = None)¶
Lorenz Conditioned Networks (LCN).
Michailidis, D., Röpke, W., Roijers, D. M., Ghebreab, S., & Santos, F. P. (2026). Scalable Multi-Objective Reinforcement Learning with Fairness Guarantees using Lorenz Dominance. Journal of Artificial Intelligence Research, 85. https://jair.org/index.php/jair/article/view/19862
LCN extends Pareto Conditioned Networks (PCN) (Reymond et al., 2022) by selecting and conditioning on commands using Lorenz dominance instead of (or interpolated with) Pareto dominance. Lorenz dominance compares returns by the cumulative sum of their sorted objectives (their Lorenz vector), which biases the learned set of policies towards more equitable trade-offs across objectives. The lcn_lambda parameter interpolates between strict Lorenz dominance and ordinary dominance computed on the sorted returns (“λ-Lorenz dominance”), trading off fairness against raw performance.
## Credits
This code extends the PCN implementation of this repository (morl_baselines.multi_policy.pcn.pcn), which is itself a refactor of the code from the authors of the PCN paper, available at: https://github.com/mathieu-reymond/pareto-conditioned-networks
Initialize LCN agent.
- Parameters:
env (Optional[gym.Env]) – Gym environment.
scaling_factor (np.ndarray) – Scaling factor for the desired return and horizon used in the model.
learning_rate (float, optional) – Learning rate. Defaults to 1e-2.
gamma (float, optional) – Discount factor. Defaults to 1.0.
batch_size (int, optional) – Batch size. Defaults to 32.
hidden_dim (int, optional) – Hidden dimension. Defaults to 64.
noise (float, optional) – Standard deviation of the noise to add to the action in the continuous action case. Defaults to 0.1.
distance_ref (str, optional) – Dominance notion used to select and condition on episodes. One of “nondominated” (strict Lorenz dominance) or “lambda_lorenz” (λ-Lorenz dominance, interpolating between Lorenz dominance and ordinary dominance on the sorted returns, controlled by lcn_lambda). Defaults to “nondominated”.
lcn_lambda (Optional[float], optional) – λ parameter used when distance_ref=”lambda_lorenz”. Controls the trade-off between fairness (Lorenz dominance) and raw performance (ordinary dominance). Required when distance_ref=”lambda_lorenz”. Defaults to None.
project_name (str, optional) – Name of the project for wandb. Defaults to “MORL-Baselines”.
experiment_name (str, optional) – Name of the experiment for wandb. Defaults to “LCN”.
wandb_entity (Optional[str], optional) – Entity for wandb. Defaults to None.
log (bool, optional) – Whether to log to wandb. Defaults to True.
seed (Optional[int], optional) – Seed for reproducibility. Defaults to None.
device (Union[th.device, str], optional) – Device to use. Defaults to “auto”.
model_class (Optional[Type[BasePCNModel]], optional) – Model class to use. Defaults to None.
- eval(obs, w=None)¶
Evaluate policy action for a given observation.
- evaluate(env, max_return, n=10)¶
Evaluate policy in the given environment.
- get_config() dict¶
Get configuration of LCN model.
- load(path: str)¶
Load LCN.
- save(filename: str = 'LCN_model', save_dir: str = 'weights')¶
Save LCN.
- set_desired_return_and_horizon(desired_return: ndarray, desired_horizon: int)¶
Set desired return and horizon for evaluation.
- train(total_timesteps: int, eval_env: Env, ref_point: ndarray, known_pareto_front: List[ndarray] | None = None, num_eval_weights_for_eval: int = 50, num_er_episodes: int = 500, num_step_episodes: int = 10, num_model_updates: int = 100, max_return: ndarray = None, max_buffer_size: int = 500, num_points_pf: int = 100, save_dir: str = 'weights', cd_threshold: float = 0.2)¶
Train LCN.
- Parameters:
total_timesteps – total number of time steps to train for
eval_env – environment for evaluation
ref_point – reference point for hypervolume calculation
known_pareto_front – Optimal pareto front for metrics calculation, if known.
num_eval_weights_for_eval (int) – Number of weights use when evaluating the Pareto front, e.g., for computing expected utility.
num_er_episodes – number of episodes to fill experience replay buffer.
num_step_episodes – number of episodes per training step
num_model_updates – number of model updates per episode
max_return – maximum return for clipping desired return. When None, this will be set to 100 for all objectives.
max_buffer_size – maximum buffer size
num_points_pf – number of points to sample from pareto front for metrics calculation
save_dir – directory to save model weights
cd_threshold – threshold for the crowding distance penalty
- update()¶
Update LCN model.