Lorenz Conditioned Networks

Lorenz Conditioned Networks (LCN) extends Pareto Conditioned Networks (PCN) by replacing the dominance criterion used to select and condition on past episodes. Where PCN uses Pareto dominance, LCN instead uses Lorenz dominance: each return’s objective values are sorted in increasing order and cumulatively summed (its Lorenz vector), and one return Lorenz-dominates another if its Lorenz vector is at least as high in every position. Lorenz dominance favors returns that distribute reward more equitably across objectives.

The lcn_lambda parameter controls λ-Lorenz dominance, an interpolation between strict Lorenz dominance and ordinary dominance computed on the sorted returns:

fv = lcn_lambda * sorted(returns) + (1 - lcn_lambda) * lorenz_vector(returns)

lcn_lambda=0 recovers strict Lorenz dominance (maximizes fairness), while values closer to 1 relax the fairness bias towards rewarding raw performance on the best-performing sorted objectives. Pick one of the two distance_ref options to choose between the two:

  • "nondominated" (default): strict Lorenz dominance.

  • "lambda_lorenz": λ-Lorenz dominance, requires lcn_lambda to be set.

For more details, see:

Michailidis, D., Röpke, W., Roijers, D. M., Ghebreab, S., & Santos, F. P. (2026). Scalable Multi-Objective Reinforcement Learning with Fairness Guarantees using Lorenz Dominance. Journal of Artificial Intelligence Research, 85. https://jair.org/index.php/jair/article/view/19862

class morl_baselines.multi_policy.lcn.lcn.LCN(env: Env | None, scaling_factor: ndarray, learning_rate: float = 0.01, gamma: float = 1.0, batch_size: int = 32, hidden_dim: int = 64, noise: float = 0.1, distance_ref: str = 'nondominated', lcn_lambda: float | None = None, project_name: str = 'MORL-Baselines', experiment_name: str = 'LCN', wandb_entity: str | None = None, log: bool = True, seed: int | None = None, device: device | str = 'auto', model_class: Type[BasePCNModel] | None = None)

Lorenz Conditioned Networks (LCN).

Michailidis, D., Röpke, W., Roijers, D. M., Ghebreab, S., & Santos, F. P. (2026). Scalable Multi-Objective Reinforcement Learning with Fairness Guarantees using Lorenz Dominance. Journal of Artificial Intelligence Research, 85. https://jair.org/index.php/jair/article/view/19862

LCN extends Pareto Conditioned Networks (PCN) (Reymond et al., 2022) by selecting and conditioning on commands using Lorenz dominance instead of (or interpolated with) Pareto dominance. Lorenz dominance compares returns by the cumulative sum of their sorted objectives (their Lorenz vector), which biases the learned set of policies towards more equitable trade-offs across objectives. The lcn_lambda parameter interpolates between strict Lorenz dominance and ordinary dominance computed on the sorted returns (“λ-Lorenz dominance”), trading off fairness against raw performance.

## Credits

This code extends the PCN implementation of this repository (morl_baselines.multi_policy.pcn.pcn), which is itself a refactor of the code from the authors of the PCN paper, available at: https://github.com/mathieu-reymond/pareto-conditioned-networks

Initialize LCN agent.

Parameters:
  • env (Optional[gym.Env]) – Gym environment.

  • scaling_factor (np.ndarray) – Scaling factor for the desired return and horizon used in the model.

  • learning_rate (float, optional) – Learning rate. Defaults to 1e-2.

  • gamma (float, optional) – Discount factor. Defaults to 1.0.

  • batch_size (int, optional) – Batch size. Defaults to 32.

  • hidden_dim (int, optional) – Hidden dimension. Defaults to 64.

  • noise (float, optional) – Standard deviation of the noise to add to the action in the continuous action case. Defaults to 0.1.

  • distance_ref (str, optional) – Dominance notion used to select and condition on episodes. One of “nondominated” (strict Lorenz dominance) or “lambda_lorenz” (λ-Lorenz dominance, interpolating between Lorenz dominance and ordinary dominance on the sorted returns, controlled by lcn_lambda). Defaults to “nondominated”.

  • lcn_lambda (Optional[float], optional) – λ parameter used when distance_ref=”lambda_lorenz”. Controls the trade-off between fairness (Lorenz dominance) and raw performance (ordinary dominance). Required when distance_ref=”lambda_lorenz”. Defaults to None.

  • project_name (str, optional) – Name of the project for wandb. Defaults to “MORL-Baselines”.

  • experiment_name (str, optional) – Name of the experiment for wandb. Defaults to “LCN”.

  • wandb_entity (Optional[str], optional) – Entity for wandb. Defaults to None.

  • log (bool, optional) – Whether to log to wandb. Defaults to True.

  • seed (Optional[int], optional) – Seed for reproducibility. Defaults to None.

  • device (Union[th.device, str], optional) – Device to use. Defaults to “auto”.

  • model_class (Optional[Type[BasePCNModel]], optional) – Model class to use. Defaults to None.

eval(obs, w=None)

Evaluate policy action for a given observation.

evaluate(env, max_return, n=10)

Evaluate policy in the given environment.

get_config() → dict

Get configuration of LCN model.

load(path: str)

Load LCN.

save(filename: str = 'LCN_model', save_dir: str = 'weights')

Save LCN.

set_desired_return_and_horizon(desired_return: ndarray, desired_horizon: int)

Set desired return and horizon for evaluation.

train(total_timesteps: int, eval_env: Env, ref_point: ndarray, known_pareto_front: List[ndarray] | None = None, num_eval_weights_for_eval: int = 50, num_er_episodes: int = 500, num_step_episodes: int = 10, num_model_updates: int = 100, max_return: ndarray = None, max_buffer_size: int = 500, num_points_pf: int = 100, save_dir: str = 'weights', cd_threshold: float = 0.2)

Train LCN.

Parameters:
  • total_timesteps – total number of time steps to train for

  • eval_env – environment for evaluation

  • ref_point – reference point for hypervolume calculation

  • known_pareto_front – Optimal pareto front for metrics calculation, if known.

  • num_eval_weights_for_eval (int) – Number of weights use when evaluating the Pareto front, e.g., for computing expected utility.

  • num_er_episodes – number of episodes to fill experience replay buffer.

  • num_step_episodes – number of episodes per training step

  • num_model_updates – number of model updates per episode

  • max_return – maximum return for clipping desired return. When None, this will be set to 100 for all objectives.

  • max_buffer_size – maximum buffer size

  • num_points_pf – number of points to sample from pareto front for metrics calculation

  • save_dir – directory to save model weights

  • cd_threshold – threshold for the crowding distance penalty

update()

Update LCN model.