Methodology and evidence

publication revision 18schema groupmixer-benchmark-publication-v1

Execution contract

Native benchmark resource classes
ClassDeadlineCPUAttempts
heavy300s8 pinned logical CPUsnative randomized: seeds 101, 202, 303, 404, 505 · seedless hosted: five fresh attempts · deterministic: one
standard60s8 pinned logical CPUsnative randomized: seeds 101, 202, 303, 404, 505 · seedless hosted: five fresh attempts · deterministic: one
  • Input preparation, solving, and output conversion count toward the deadline.
  • Native attempts run in fresh processes with an 8 GiB process limit.
  • Returned schedules are independently validated without repair.
  • Hosted and native runtimes are reported but not compared.

Status vocabulary

No candidate, timeout, error, and infeasibility are separate outcomes.

Benchmark publication status definitions
OptimalA zero lower bound or declared target was attained.
Optimal, provenThe solver returned a matching bound and verified proof evidence.
FeasibleAt least one canonically valid schedule was returned.
Infeasible, provenThe exact hard-feasibility problem was proven infeasible.
No candidate timeoutThe declared search envelope ended without a candidate.
ErrorThe process or evidence pipeline failed; no timeout or infeasibility is inferred.
UnsupportedThe tool cannot represent the case or exceeds an observed product limit.
UnavailableA runnable licensed environment was not available.
Invalid outputOutput was returned but failed canonical validation.
Pending verificationThe workflow or capability has not entered publication evidence.
Not runNo execution was admitted for this row.

Configuration rules

Allowed

  • Case-specific formulations and symmetry.
  • Disclosed runtime construction and staged search.
  • The strongest frozen end-user setting admitted within the case resource envelope.

Forbidden

  • Embedded schedules, answers, or cross-solver hints.
  • Proxy cases, weakened semantics, or repaired output.
  • Selection by a metric outside the declared objective.

Authorship disclosure

GroupMixer authors this benchmark and the published OR-Tools formulations. Returned schedules are evaluated against the same canonical scenario rules. This is not an independent third-party benchmark.

Publication files

SHA256SUMS
Publication evidence files with sizes and SHA-256 checksums
FileSizeSHA-256
publication.json314.7 KiBsha256:ac8d008e34560c89f65ae3738498593e6a7b99a3d1159a084f4f4a828880b6bb
attempt-aggregate.json482.2 KiBsha256:7f8c3568e0ef956324208effb9b1fa6f7d4875a35663069e22e83396529726e0
publication.schema.json9.0 KiBsha256:589eda3dfa619da5352ff62acf0e2a16f30b5fc1ea469c3e59a06719d5aba4ad
evidence.tar.gz4.5 MiBsha256:e6fad18840e34bfc08893581cd2726638289a6f70320df851d97021f5d1c36cc

Reproduce and verify

cd webapp && npm run benchmark:data:check
jq '.summaries[]' public/benchmarks/public-group-generator-v1/revision-18/attempt-aggregate.json
cd public/benchmarks/public-group-generator-v1/revision-18 && sha256sum -c SHA256SUMS

Comparator evidence revisions: r1 (159 rows) · r2 (14 rows) · r3 (1 rows) · r4 (1 rows) · r5 (2 rows) · r6 (5 rows). Dry and superseded evidence is excluded.