![]() |
Ansel 0.0
A darktable fork - bloat + design vision
|
#include "develop/pipe_cache_policy.h"#include <stdarg.h>#include <stddef.h>#include <setjmp.h>#include <stdint.h>#include <cmocka.h>
Include dependency graph for test_pipe_cache_policy.c:Go to the source code of this file.
A CPU-only node in the middle publishes its producer and NOTHING further. This is the shape the transitive version was written for (colorout -> rawoverexposed -> dither): the producer must publish, and the nodes before it must not have to.
Definition at line 238 of file test_pipe_cache_policy.c.
References _gpu_node_with_no_needs(), _walk(), FALSE, i, L, state, dt_dev_pipe_cache_policy_inputs_t::supports_opencl, TRUE, and void().
Referenced by main().
A CPU-only node still makes the node before it publish, which is the whole point of the one hop. This is the case the transitive version existed to protect and the lean one must keep protecting.
Definition at line 112 of file test_pipe_cache_policy.c.
References dt_dev_pipe_cache_policy_decide(), FALSE, L, state, and void().
Referenced by main().
THE REGRESSION, restated. An enabled, GPU-capable node that needs no host input of its own must NOT hand a downstream requirement further upstream.
This assertion is the opposite of the one it replaces, and deliberately so. The requirement being carried is "the node that CONSUMES my output reads it from RAM". That is a fact about one edge of the graph; the node before me publishes to a consumer of its own and knows nothing about mine. Carrying it made the flag monotone, and since the seal seeds the walk with TRUE for the displayed final output, EVERY enabled node inherited it: measured on a painting stroke, 152 MB and 68.1 ms of a 113.8 ms frame spent copying module outputs to RAM, of which only the last node's 11.7 MB was ever read from RAM by anything.
The defect the transitive version was written for is real but lives elsewhere: a node whose requirement turns back ON can be handed a cacheline whose host copy was never refreshed while the requirement was off. _seal_opencl_cache_policy() invalidates the line on that edge. Keeping every host copy fresh forever also prevented it, at the cost above.
Definition at line 93 of file test_pipe_cache_policy.c.
References _gpu_node_with_no_needs(), dt_dev_pipe_cache_policy_decide(), L, state, TRUE, and void().
Referenced by main().
The lean rule, stated whole: a GPU node nothing on the host reads must cache NOTHING.
Caching an OpenCL module's output to RAM is justified by exactly two things – the module being expensive enough that code or the user pinned it, and its output being read from RAM by a histogram, a colour picker or a Cairo surface. A node with neither, whose consumer is also on the GPU, hands its output to nobody on the CPU and must keep it on the device.
Definition at line 130 of file test_pipe_cache_policy.c.
References _gpu_node_with_no_needs(), dt_dev_pipe_cache_policy_decide(), FALSE, L, state, TRUE, and void().
Referenced by main().
With nothing displayed and nothing on the CPU, the chain caches nothing at all.
Definition at line 259 of file test_pipe_cache_policy.c.
References _gpu_node_with_no_needs(), _walk(), FALSE, i, L, state, and void().
Referenced by main().
A node that cannot run on the GPU at all: it produces host data by construction.
Definition at line 49 of file test_pipe_cache_policy.c.
References dt_dev_pipe_cache_policy_decide(), FALSE, L, state, dt_dev_pipe_cache_policy_inputs_t::supports_opencl, and void().
Referenced by main().
Each of the five own-input reasons must raise the upstream requirement by itself.
Definition at line 141 of file test_pipe_cache_policy.c.
References _gpu_node_with_no_needs(), dt_dev_pipe_cache_policy_decide(), FALSE, i, L, state, TRUE, and void().
Referenced by main().
Each of the four own-output reasons must raise this node's own requirement by itself.
Definition at line 163 of file test_pipe_cache_policy.c.
References _gpu_node_with_no_needs(), dt_dev_pipe_cache_policy_decide(), FALSE, i, L, state, TRUE, and void().
Referenced by main().
Definition at line 64 of file test_pipe_cache_policy.c.
References _gpu_node_with_no_needs(), dt_dev_pipe_cache_policy_decide(), FALSE, L, state, TRUE, and void().
Referenced by main().
|
static |
Which pipeline nodes must keep a host-RAM copy of their output.
These tests exist because nothing else can see this. The per-piece cache_output_on_ram flag changes no exported pixel – the export path drives the pipe directly and never reaches the seal that computes it – changes no hash, and produces no log unless someone is already looking. When it was wrong, the symptom was a downstream module reading stale host bytes from a rekeyed cacheline, only with OpenCL enabled, and it was found by dumping GPU buffers.
The propagation test below is that bug, pinned. A GPU-capable node with no reason of its own to want host data.
Definition at line 41 of file test_pipe_cache_policy.c.
References dt_dev_pipe_cache_policy_inputs_t::supports_opencl, and TRUE.
Referenced by _a_cpu_node_publishes_its_producer_and_no_further(), _a_gpu_node_does_not_relay_its_consumers_requirement(), _a_gpu_node_nothing_reads_from_ram_caches_nothing(), _a_pure_gpu_chain_nobody_reads_caches_nothing(), _each_own_input_reason_raises_upstream(), _each_own_output_reason_raises_own(), _gpu_node_requires_nothing_on_its_own(), _null_upstream_pointer_is_allowed(), and _only_the_displayed_output_is_cached_in_a_gpu_chain().
A NULL out-param is allowed: the last node in the walk has nobody upstream to tell.
Definition at line 186 of file test_pipe_cache_policy.c.
References _gpu_node_with_no_needs(), dt_dev_pipe_cache_policy_inputs_t::color_picker_on, dt_dev_pipe_cache_policy_decide(), FALSE, L, state, TRUE, and void().
Referenced by main().
The darkroom case, and the one that was costing 152 MB a frame: every module on the GPU, the last one's output displayed from host memory by Cairo. Exactly one node may cache.
Definition at line 218 of file test_pipe_cache_policy.c.
References _gpu_node_with_no_needs(), _walk(), FALSE, i, L, state, TRUE, and void().
Referenced by main().
|
static |
Walk a chain of nodes the way _seal_opencl_cache_policy() does, last to first.
The policy is a per-node decision, but what it MEANS is a property of the chain: the seal threads one flag backwards through every enabled node. The per-node tests above cannot see that, and the defect they missed lived entirely in the composition – a requirement that was correct for one node became, by being relayed, a requirement for every node before it. So these run the walk.
out[] receives each node's cache_output_on_ram, indexed as nodes[] is.
Definition at line 206 of file test_pipe_cache_policy.c.
References dt_dev_pipe_cache_policy_decide(), i, L, and out.
Referenced by _a_cpu_node_publishes_its_producer_and_no_further(), _a_pure_gpu_chain_nobody_reads_caches_nothing(), and _only_the_displayed_output_is_cached_in_a_gpu_chain().
| int main | ( | void | ) |
Definition at line 271 of file test_pipe_cache_policy.c.
References _a_cpu_node_publishes_its_producer_and_no_further(), _a_cpu_node_still_makes_its_producer_publish(), _a_gpu_node_does_not_relay_its_consumers_requirement(), _a_gpu_node_nothing_reads_from_ram_caches_nothing(), _a_pure_gpu_chain_nobody_reads_caches_nothing(), _cpu_only_node_needs_its_input_on_host(), _each_own_input_reason_raises_upstream(), _each_own_output_reason_raises_own(), _gpu_node_requires_nothing_on_its_own(), _null_upstream_pointer_is_allowed(), _only_the_displayed_output_is_cached_in_a_gpu_chain(), and L.