Writing Evaluation Callbacks
During optimization ropt decides which variable vectors need values for the
objectives and optional nonlinear-constraints. A compute step does not compute
these values itself — it delegates to an
Evaluator instance that you supply. This
page describes the evaluators ropt provides and how to write the evaluation
code they wrap.
The evaluators
| Evaluator | Interface |
|---|---|
BatchEvaluator |
Batch: f(variables_2d, context) → EvaluationBatchResult. |
FunctionEvaluator |
Per-row: f(variables_1d, context) → EvaluationFunctionResult. |
CachedEvaluator |
Wraps another evaluator, caching results by variable vector. |
ParallelEvaluator |
Parallel evaluation via an Executor — see Parallel Evaluation. |
The first three run synchronously in the calling thread;
ParallelEvaluator dispatches
work to an Executor and is described in
Parallel Evaluation.
Evaluators are not safe for concurrent use
An evaluator raises a
WorkflowError if two threads execute its
eval method at the same time. Serial reuse is allowed: a single evaluator
instance may be shared by several compute steps that run one after another,
even on different threads (for example, reusing one FunctionEvaluator
across nested inner optimizations to keep batch ids counting). Do not
share one evaluator across steps that run in parallel; give each parallel
step its own evaluator. For the constraints on where each layer of a nested
workflow may run, see
Nested workflows and process boundaries.
Note that the parallelism of
ParallelEvaluator happens
below eval — it dispatches tasks to an executor, so its own eval is
still called on a single thread. An evaluator cannot be transferred to
another process: it holds a lock, so serializing one fails and the
submission is refused with an
ExecutionError.
You write evaluation code for two of them.
BatchEvaluator takes a callback
that receives the full 2-D batch of variable vectors;
FunctionEvaluator wraps a
simpler function called once per row. Both are covered below, followed by
CachedEvaluator, which wraps
another evaluator to reuse previously computed results.
Writing a batch callback
Evaluation callbacks must adhere to the
EvaluationBatchCallback protocol.
For instance an evaluator function should look like this:
from numpy.typing import NDArray
import numpy as np
from ropt.evaluation import EvaluationBatchContext, EvaluationBatchResult
def my_evaluator(
variables: NDArray[np.float64],
context: EvaluationBatchContext,
) -> EvaluationBatchResult:
...
variableshas shape(n_rows, n_variables). Each row is a separate variable vector to evaluate.contextcarries per-row metadata (see below) plus the immutableEnOptContextfor the run.- The return value should be an
EvaluationBatchResultobject that packages objective values (and optional constraint values, metadata, and per-row error indicators).
One advantage of this approach is that the callback receives all variable vectors at once as a 2-D NumPy array. This makes it possible to exploit NumPy's vectorized operations to evaluate all rows in a single pass, avoiding explicit Python loops and achieving better performance.
What is in EvaluationBatchContext
The EvaluationBatchContext dataclass exposes:
| Field | Meaning |
|---|---|
context |
The full EnOptContext (read-only). |
active |
A boolean array indicating which rows actually need evaluation. |
realizations |
Integer realization index for each row. |
perturbations |
Integer perturbation index per row, or -1 for unperturbed rows. None if no perturbations are used. |
Use realizations to pick the right per-realization model parameters (an
uncertainty draw, a different simulation deck, etc.). Use active to skip rows
that are not needed: utility methods
get_active_evaluations
and
insert_inactive_results
help filter the input and re-expand the output.
Returning results
EvaluationBatchResult stores:
| Field | Meaning |
|---|---|
objectives |
Objective values, shape (n_rows, n_objectives). |
constraints |
Optional constraint values, shape (n_rows, n_nonlinear_constraints). |
metadata |
Optional per-row metadata dict; not used by ropt. |
batch_id |
Batch label (default 0). |
constraints is required when nonlinear_constraints is configured in the
problem. metadata is stored verbatim on the resulting
Results object and is useful for linking results back
to the input vectors that produced them. batch_id defaults to 0; all
results will carry this label unless you set it yourself. For
auto-incrementing IDs pass a
BatchIdCounter (or any
Callable[[], int]) to the batch_id_callback argument of
FunctionEvaluator or
ParallelEvaluator; for raw
BatchEvaluator callbacks set it
yourself.
Inactive rows (where active is False) should have their result values set
to zero. Rows where an evaluation failed should be set to np.nan (see
Handling partial failures below).
Returning constraints
def evaluator(variables, context):
obj = ... # shape (n_rows, n_objectives)
con = ... # shape (n_rows, n_nonlinear_constraints)
return EvaluationBatchResult(objectives=obj, constraints=con)
lower_bounds / upper_bounds
declared in
NonlinearConstraintsConfig.
Handling partial failures
If your evaluator cannot compute a given objective, set the corresponding entry
in the objectives field to np.nan. ropt treats NaN rows as failed
evaluations; the realization_min_success and
perturbation_min_success settings determine
whether the optimization can recover. For example:
def evaluator(variables, context):
n_rows, n_obj = variables.shape[0], 1
obj = np.full((n_rows, n_obj), np.nan)
for row in range(n_rows):
try:
obj[row, 0] = simulate(variables[row])
except SimulationError:
pass # leave NaN
return EvaluationBatchResult(objectives=obj)
Combined with the realization_min_success field of
RealizationsConfig, this allows the
optimization to continue as long as enough realizations succeed.
Using FunctionEvaluator
When your evaluation function naturally works on a single variable vector at a
time — for instance when it calls an external simulator once per realization —
the FunctionEvaluator offers a
simpler alternative. Instead of receiving the full 2-D batch and managing the
loop yourself, you write a function that takes a single 1-D variable vector and
returns the objective (and optional constraint) values for that row. The
FunctionEvaluator handles the batching, the active-row filtering, and the
assembly of the final
EvaluationBatchResult.
A function passed to FunctionEvaluator must follow the
EvaluationFunctionCallback protocol:
from numpy.typing import NDArray
import numpy as np
from ropt.components.evaluators import (
EvaluationFunctionContext,
EvaluationFunctionResult,
)
def my_function(
variables: NDArray[np.float64],
context: EvaluationFunctionContext,
) -> EvaluationFunctionResult:
...
variablesis a 1-D array for a single evaluation row.-
contextis anEvaluationFunctionContextdataclass identifying the evaluation. It exposes:Field Meaning realizationInteger realization index for this row. perturbationPerturbation index, or -1when unperturbed.batch_idInteger identifying the current evaluation batch. eval_idxRow index within the batch. metadataThe metadata the run was started with, or None.metadatais the dictionary passed to the compute step'srunmethod, and is the channel for handing your own per-run data (a run ID, a case name, a path) to the evaluation function. Every evaluation of the run sees the same dictionary, so treat it as read-only; for a process or HPC executor its contents must be picklable. -
The return value is an
EvaluationFunctionResultdataclass with the following fields:Field Meaning objectivesScalar or 1-D array of length n_objectives.constraintsOptional scalar or 1-D array of length n_nonlinear_constraints.metadataOptional dict[str, Any]; stored verbatim in the batch result.Each
metadataentry is forwarded intoEvaluationBatchResult.metadatafor the corresponding row.A key need not be set by every row. Rows that do not set it get
np.nanfor numeric values andNoneotherwise, so a numeric column with missing rows is widened tofloat64— numpy has no integer NaN. A column set by every row keeps its natural dtype. Mixing strings and non-strings under one key raises aValueError.A value may be a 1-D array rather than a scalar, which gives the key an extra user-defined axis named after the key. Every row must then return the same number of entries, and the key may not be named after an
AxisNamevalue orbatch_id.
Wrap the function in a
FunctionEvaluator and give the
evaluator to a compute step (see Optimization Workflows):
from ropt.components.evaluators import FunctionEvaluator
evaluator = FunctionEvaluator(function=my_function)
Using CachedEvaluator
CachedEvaluator wraps another
evaluator with result caching. It retrieves previously computed function results
from EventHandler instances specified as sources — typically a
HistoryHandler or ResultsHandler. For each variable vector and realization,
if a matching cached result is found, the cached objectives and constraints are
reused without calling the wrapped evaluator. Only uncached evaluations are
forwarded to the underlying evaluator.
Cache matching works as follows: for each requested variable vector and
realization, the evaluator searches through the "results" stored by its
sources. A match is found when the variables are equal (within floating-point
tolerance) and the realization matches. If realization names are configured,
they are used for matching (allowing cache hits across different optimization
runs with the same realization names). Otherwise, realization indices are used.
If some but not all evaluations are found in cache, the cached ones are marked as inactive and only the missing evaluations are delegated to the wrapped evaluator. The final combined result contains both cached and newly computed values.
Sources can be added dynamically with add_sources().
To record which evaluations were served from cache, pass a hits_key string
at construction time. When set, the returned
EvaluationBatchResult will contain
a boolean NumPy array in its metadata dictionary under that key —
True for evaluations that came from the cache, False for those that were
freshly computed.
The eval_cached() method is available for derived classes that need access to
which evaluations were cache hits — it returns both the
EvaluationBatchResult and a
dictionary mapping evaluation indices to their cached
FunctionResults.
Using ParallelEvaluator
The evaluators above run each function call sequentially in the current thread.
For parallel evaluation — whether via worker threads, separate processes, or an
HPC cluster — use
ParallelEvaluator. See
Parallel Evaluation for the evaluator and the available
executors.
Reusing objectives and constraints
When defining multiple objectives, you may need to reuse the same underlying computation. For example, a total objective could consist of the mean of the realizations plus their standard deviation. Rather than evaluating all realizations twice, compute them once and return the values for both objectives from a single evaluator call.
Where to next
- Read the results: Working with Results.
- Run evaluations in parallel, in processes, or on a cluster: Parallel Evaluation.
- See it in action: Building a Workflow.