Files
OrcaSlicer/tests/AGENTS.md
T
packerlschupfer 9321f24959 CLI: --strict, and a warnings array in result.json (#14601)
# Description

Add `--strict` for CI and scripted pipelines, and a structured
`warnings`
array in `result.json`.

## `--strict`

A NON_CRITICAL slicing warning is logged and the slice succeeds: return
code
`0`, G-code written. That suits interactive use, but a pipeline then
ships a
slice with a warning nobody saw. With `--strict`, such a warning fails
the run
with `CLI_SLICING_ERROR` before the G-code is exported. Without the
flag,
nothing changes.

In FFF the warning that reaches this path is "support needed but
disabled"
(`PrintObject::generate_support_material`). `--no-check` skips that
check, so
`--strict --no-check` is rejected with `CLI_INVALID_PARAMS`.

`--strict` is read before any work, so it doesn't depend on argument
order and
`result.json` reports it for early failures as well.

## `result.json`

Two new top-level fields:

- `warnings`: `[{"class", ...details}]`. One class is wired:
`slicing_warning_non_critical` with `plate_id` and `text`, recorded
whenever
such a warning fires, with or without `--strict`. The array also fills
on
  runs that succeed, so `return_code` stays the verdict.
- `strict_mode`: whether `--strict` was on.

`record_exit_reson` writes `result.json` on Linux only, so both fields
exist
only there. The non-zero exit works on every platform.

## Tests

- `tests/fff_print/test_support_material.cpp` (all platforms): an
overhang
sliced with support off raises the NON_CRITICAL support-needed status,
and
  the no-check flag suppresses it.
- `tests/cli/test_cli_strict.sh` (Linux only): runs `orca-slicer`
without
flags, with `--strict`, and with `--strict --no-check`, and checks the
shell
status and `result.json` of each. It runs the built binary, so it
carries the
`RequiresApp` label, which `scripts/run_unit_tests.sh` excludes because
the
  unit-test job only receives `build/tests`. Run it with
  `ctest --test-dir build/tests -C Release -L RequiresApp`.
- CI: `unit_tests.yml` now passes `Release` on Linux too.
`build_linux.sh`
configures Ninja Multi-Config, and without a config ctest drops the
labels of
plain `add_test()` tests, so this test ran as "Not Run" instead of being
excluded. The docs that assumed Linux was single-config are corrected
too.

Built and run locally on Linux (GCC 14) on current `main`: both tests
pass,
and the touched files compile clean under Clang with `-Werror`.
2026-09-16 12:54:48 +08:00

5.0 KiB

Test suite rules

Rules for writing tests under tests/. CATCH2.md is the Catch2 reference. Building and running the suites is covered on the wiki, at https://www.orcaslicer.com/wiki/developer_reference/how_to_test.html.

The suites

  • libslic3r: the core library. Geometry, meshes, file formats, config and presets, Clipper, algorithms, data structures.
  • fff_print: the FFF slicing pipeline, from a Model plus config through Print and PrintObject to emitted G-code.
  • sla_print: SLA support-tree and pad geometry, support-point generation, raycast.
  • libnest2d: 2D nesting and packing.
  • slic3rutils: the Python plugin system and its slicing-pipeline bindings.
  • filament_group: filament-to-extruder grouping, checked against golden files.
  • cli: end-to-end runs of the built orca-slicer binary, Linux only. These tests carry the RequiresApp label, which the CI unit-test job excludes because it receives only build/tests; run them with ctest --test-dir build/tests -C Release -L RequiresApp.

Building and running

Tests are off by default, so the build has to be told to include them.

  • Windows: build_release_vs.bat tests, then ctest --test-dir build/tests -C Release
  • macOS: ./build_release_macos.sh -s -a arm64 -T, which builds and runs them
  • Linux: ./build_linux.sh -t, then ctest --test-dir build/tests -C Release

Rebuild a single suite with cmake --build build --config Release --target <suite>_tests. Visual Studio, Xcode and the Ninja Multi-Config generator that build_linux.sh uses are all multi-configuration, so ctest needs -C on every platform; without it, tests registered with plain add_test() lose their labels and report "Not Run".

Where a test goes

  • Pick the suite by the production code the test exercises, not by how the test is written.
  • A property of a class that holds with no Print involved belongs in libslic3r. Behavior that depends on print settings, or produces or consumes G-code or slicing state, belongs in fff_print.
  • One file per subsystem, named test_<subsystem>.cpp. It owns every test for that subsystem, whether the test reads in-memory state or generated output.
  • When you add a file, list it in that suite's CMakeLists.txt in the same change.

Use the existing helpers

Check these before writing your own setup or output-parsing code.

  • tests/test_utils.hpp is shared by every suite. load_model() loads a mesh from tests/data/, and ScopedTemporaryFile gives a temp path that removes itself.
  • fff_print/test_helpers.hpp builds and slices a Print and parses the emitted G-code. Read it before writing an fff_print test rather than assembling a Print by hand.
  • The other suites have their own: sla_print/sla_test_utils.hpp, libnest2d/libnest2d_test_utils.hpp, slic3rutils/plugin_test_utils.hpp, filament_group/fg_test_utils.hpp. libslic3r has none and uses the shared header.
  • Test data lives in tests/data/ and is reached through the TEST_DATA_DIR define. Wrap it in std::string(...) before joining a path onto it.

Writing the test

  • Name the test case as a plain behavioral sentence in the present tense. No Subsystem: prefix.
  • Tag it with the subsystem it covers, matching the file, in PascalCase. That tag is what people filter on, so every test needs one.
  • Add further tags where they help: a narrower one to slice a large file ([Rotcalip], [Placer]), a shared one for something spanning files ([Python], [H2C], [Regression]), or [NotWorking] / [.] to disable or hide a test. Say why in a comment if you disable or hide.
  • Prefer a flat TEST_CASE per behavior, with GENERATE for parameterized cases. Reserve SCENARIO / GIVEN / WHEN / THEN for genuine shared setup that branches into a few close variations.
  • Set the config keys your test depends on, and derive the expected values from what you set. A 20mm cube sliced at layer_height 2 is 10 layers, and the test should state both parts. If a number in your assertion comes from a key you never set, the test is also testing that default.
  • Assert the defining property, not an incidental value. "Skirt present" or "at least 2 brim loops" survives a refactor; exact coordinates and byte counts do not.
  • Name a regression test for the behavior it protects, never for an issue or PR number.
  • When asserting on G-code, match the meaningful token such as ; skirt rather than whole lines, whitespace or comment wording. Depend on ordering only when ordering is the contract.

Catch2 rules that cause real breakage

  • Never reuse a SECTION name inside a loop. Use DYNAMIC_SECTION so each iteration is unique.
  • Never assert from a spawned thread. Catch2 assertions are not thread-safe. Collect results in the thread and assert on the main thread.
  • Never combine conditions with && or || inside one assertion. Split them so Catch2 can print both operands on failure.
  • Compare floats with WithinAbs or WithinRel, never ==. Prefer these over Approx in new tests.
  • Keep tests self-contained: no shared state, green under --order rand.