![]() |
Ansel 0.0
A darktable fork - bloat + design vision
|
First checked 2026-09-29. This file was mechanically checked against
8f4638a04eon 2026-09-29 — everyfile:linecitation resolved, every backticked symbol looked up in the tree, every OPEN/planned status claim tested, and every gate or baseline number it quotes compared withtools/check_module_boundaries.shandtools/include_baseline.txt. No per-claim semantic read was done: a citation that resolves can still describe the wrong thing, so this is a floor, not a verification. Two numbers corrected: layering violations are 183, not 217, and one function was renamed when the database module was sealed (dt_history_repository_foreach_row()).
Darktable was not modular, as the /src dependency graph shows: everything is wired to the GUI, IOP "modules" are all aware of everything. Modifying anything somewhere typically broke something unexpected somewhere else.
The end goal is a backend that does not know a GUI exists, and a frontend thin enough that replacing GTK with Qt is a sizeable job rather than a rewrite. Everything below serves that.
Where the code stands is in § Module map and the rules that keep it there are in § The rules. Read those two before changing anything structural. The rest of this document describes the architecture those rules protect.
We should here make distinctions between:
IOP (Image OPerations) modules are more like plugins: it's actually how they are referred to in many early files. They define both a pipeline node (aka pixel filtering code) and a GUI widget in darkroom. The early design shows initial intent of making them re-orderable in the pipelpine, and to allow third-party plugins. As such, the core was designed to be unaware of IOP internals.
As that initial project seemed to be abandonned, IOP modules became less and less enclosed from the core, which allowed some lazy mixes and confusions between what belongs to the scope of the pipeline, and what belongs to the scope of modules. The pipeline output profile can therefore be retrieved from the colorout module or from the pipeline data. color calibration reads the input profile from colorin.
IOP modules also commit their parameter history directly using `darktable.develop` global data, instead of using their private link `(dt_iop_module_t *)->dev`, which is documented in an old comment to be the only thread-safe way of doing it.
The editing history is the snapshot of parameters for each IOP. It gets saved to the database. It gets read and flattened to copy parameters to pipeline nodes.
The development is an hybrid thing joining history with pipeline. This object will take care of reading the image cache to grab the buffer, reading the database history, initing a pipeline, and starting a new computing job.
flowchart TD;
subgraph Modules
M((IOP))
G{GUI}
F{Defaults}
end
subgraph Develoment
P((Pipeline))
H{History}
end
P --update--> M
P --read--> H
G --read---> C
G --write---> H
G --read---> H
M --fetch--> C
M --push--> C
H --read--> D
H --write--> D
D[Database]
C[Pipe cache]
I[Image cache]
A[Mipmap cache]
I --read---> D
A --write--> I
F --read--> I
H --read--> F
L{Lighttable} --write--> I
gui_init(), gui_update(), gui_changed(), gui_cleanup() and `module->params`).piece object holds a reference to the base `module` which should never be used for write operations from within pipelines because it lives in GUI thread,We control the pipeline through the history, not through the GUI. So color-pickers work by having modules GUI wait for pipeline output, sample and write new parameters into history. Modules and pipeline are opaque to the history.
Caches are high-level manager that host a list of cachelines (whether as GList or as GHashtable). They all have an high-level dt_pthread_mutex_lock_t lock which is held for a minimal amount of time while changing the size of the cachelines list (add/remove cachelines) or looking for cachelines (in which case we don't want the list to change size in the process). Once a cacheline is found or created, it has its own read/write dt_pthread_rw_lock_t lock : writing is slow because it waits for all reading threads to complete, and only one thread can write at any given time, but reading is fast because it only waits for writing to complete and several threads can concurrently read. Write locks must therefore be held for a minimal amount of time to not freeze the whole application unnecessarily.
As a general rule, thread locks should be promoted to read/write locks when it is safe for several threads to concurrently read from the same buffers.
The database is handled through SQLite3, which is not thread-safe in itself, so writing the same data for the same image from concurrent threads leads to undefined behaviour. We achieve thread-safety through locking in write mode the image cache entry of the manipulated image whenever performing read/write DB operations. But that needs to be carefully implemented in each layer (development history, tagging, rating, colorlabels, metadata, etc.) and I can't guarantee I didn't forget some spots.
At some point, we will need to choose between:
In any case, direct database reads/writes should be removed entirely from all the application because ensuring thread safety can only happen if they all happen from a central place, exposed to modules and core through a public API that prevents messing with the internals.
The image cache holds the basic representation of an image, fetched from the database : filename, path, filmroll ID, ratings, color labels, size, date/time, flags (RAW type, history inited status, etc.), orientation. It is therefore a symbolic representation of the image.
The one thing that it does not grab from database is the image buffer descriptor, which is inited by the image codecs when the actual image file is loaded into the mipmap cache. This design is brittle and should be fixed (by saving dt_image_t.dsc to database too), because the current design means that we can never assume having a dt_image_t object for a particular image ID means this object is completely defined when we read it, and it will never be if the image is not present on disk at the moment (but hosted on unavailable NAS or unplugged harddrive). The function dt_dev_ensure_image_storage() will chain loading both the image cache and the mipmap cache for a particular image as to guarantee that.
IOP modules' default_params() method, that initialize the internal parameters when building fresh histories for new pictures, need complete dt_image_t object, including the dt_image_t.dsc buffer descriptor, which is used in particular to force demosaicing on RAW images or disable it on non-RAW images.
The two-phase (provisional → resolved) lifecycle that lets the application classify an image before it is decoded — and the canonical, non-overlapping type API (dt_image_pipe_class() and the orthogonal dt_image_needs_* predicates) that replaces the overlapping dt_image_is_* heuristics — are documented separately in image-type-detection.md.
The GUI layer interfacing with dt_image_t object for user read/write is the lighttable, through the thumbnails.c and thumbtable.c libraries.
The mipmap cache loads images from the disk, decodes them through codecs from src/imageio/imageio_* and stores them in RAM. It is used for cached thumbnails as well as input RAW images. As mentionned above, the codecs update the dt_image_t.dsc buffer descriptor.
The mipmap cache is used when initing histories in dt_dev_load_image() and in thumbnail.c when fetching the base image to display in thumbnails.
IOP modules read their pixel input from the pipeline cache, and write their pixel output into it. They also publish side-band raster masks there under dedicated keys. Then, the GUI parts that need image surfaces read them directly from the pipeline cache too, as well as all the histograms and color-picker samplings.
Regular image cachelines are indexed by the dt_dev_pixelpipe_iop_t.global_hash of their pipeline node piece object. To find the pipeline node connected to a certain module in a certain pipeline, use the function dt_dev_pixelpipe_get_module_piece(). From there, two different functions can fetch the data buffer connected to the cacheline associated to that piece:
dt_dev_pixelpipe_cache_peek_gui() will either return the cacheline and data buffer if available, or queue a cache-wait and request a partial pipeline recompute (up to this module) to create it if missing,dt_dev_pixelpipe_cache_peek() will return the cacheline and data buffer if available, or nothing.Note that a module's input is the output of the previous enabled module (dt_dev_pixelpipe_get_prev_enabled_piece()), so GUI code that needs a module's input must fetch the previous piece's cacheline, not the module's own. The full mechanics of the GUI fetch — the cache-wait manager, the DT_SIGNAL_CACHELINE_READY retry protocol, the partial-recompute request, and the pitfalls of using raw dt_dev_pixelpipe_cache_peek() from GUI — are documented separately in pipeline-cache.md.
Raster masks use the same global cache but a different identity and lifecycle. dt_dev_pixelpipe_raster_mask_hash() derives a dedicated key from the provider's global_mask_hash, a namespace tag and the provider-local mask ID. The provider publishes one immutable single-channel float buffer; consumers copy it under a read lock and apply downstream distort_mask() callbacks to their private working copy. The pipe retains the mask cachelines required by its current graph in dt_dev_pixelpipe_t.raster_mask_hashes, so preview, full and export pipelines can reuse masks even when the provider image is an exact cache hit. A rare mask-only eviction in an interactive pipe triggers one targeted DT_DEV_PIPE_REENTRY pass from the provider, without rebuilding the pipeline nodes or flushing upstream image states. Export enumerates the mask IDs declared by each provider and reads the same dedicated cachelines directly. See the Raster masks are dedicated side-band cachelines section in pipeline-cache.md.
This new pipeline architecture allows for fully asynchronous pipelines, where IOP modules can grab their input from the output of any arbitrary module (even non-sequential), pipelines can have parallel branches, and we can easily run partial pipelines (starting or ending at any arbitrary node).
The protocol creation of the dt_dev_pixelpipe_iop_t.global_hash hash is detailed in the sequence of functions that create it:
dt_iop_compute_blendop_hash(), writing dt_iop_module_t.blendop_hash from internal blending and masking parameters,dt_iop_compute_module_hash(), writing dt_iop_module_t.hash from internal parameters,dt_dev_history_item_update_from_params() and dt_dev_read_history_ext(), initing dt_dev_history_item_t.hash with dt_iop_module_t.hash,dt_dev_set_history_hash() aggregating all the history items dt_dev_history_item_t.hash over the whole history stack to produce one single integrity checksum for the current history. Note: that checksum is used by the darkroom pipelines to trigger new recomputations if needed (aka checksum changed)dt_iop_commit_params(), writing the piece-wise (pipeline node) dt_dev_pixelpipe_iop_t.hash and dt_dev_pixelpipe_iop_t.blendop_hash aggregating both the static hashes dt_iop_module_t.hash and dt_iop_module_t.blendop_hash, with dynamic, runtime-defined parameters inited at commit_params() time.dt_pixelpipe_get_global_hash(), writing the dt_dev_pixelpipe_iop_t.global_hash from dt_dev_pixelpipe_iop_t.hash and the dt_dev_pixelpipe_t states (mask preview mode, cache bypass modes, etc.), finally used as cacheline ID.So the application doesn't have explicit pipeline rendering triggers anymore or invalidation flags, checksums/hashes track the internal states of all data structures across the software, and GUI updating functions as well as pipeline rendering functions keep track of the checksum of the objects they are interested in, use them to fetch the associated pixel buffers from the pipeline cache, or to trigger internal updates/renderings if they are not available on the cache.
This removes a lot of pressure on multi-threading synchronization because threads can live entirely in their own timeline without having to wait each other or start each other. We just declare data states through checksums and let each thread decide what it should do with it.
Also, when a new pipe cacheline is written, it raises the signal DT_SIGNAL_CACHELINE_READY with the hash of the cacheline, which means that all GUI places waiting for a particular buffer rendering can connect on this signal and immediately refresh their internal state without having to wait for a full pipeline to complete. GUI consumers should not subscribe to this signal directly: the shared cache-wait manager (dt_dev_pixelpipe_cache_peek_gui() + dt_dev_pixelpipe_cache_wait_t) already centralizes the subscription, the deduplication of pending requests, the busy-cursor feedback and the resume callbacks. See pipeline-cache.md.
As stated above, reads are "fast" and writes are "slow", insofar as concurrent reads are thread-safe while writes need to wait for each other and for reads to finish. Moreover, it usually doesn't make sense that several threads would write the same data : threads are specialized for one single task. Beyond developer OCD, this allows to track reliably the lifecycle of data and expose the minimum amount of API outside of each library, while keeping data management as centralized and as private as possible.
Some places violate this principle, and shouldn't be used as precedents to justify and legitimate further extension of bad design, but should be taken as TODO to fix the architecture and its APIs.
Here is a complete image lifecycle, assuming it is already imported into database:
dt_image_t object from database to image cache through dt_image_cache_get(),dt_image_t.dsc structure through dt_mipmap_cache_get(),dt_develop_t.history through dt_dev_read_history_ext() which, internally:dt_dev_init_default_history(), with module default parameters, auto-presets and mandatory modules for the image type,dt_history_repository_foreach_row() (src/database/history_repository.h:118; it was dt_history_db_foreach_history_row() before the database module was sealed),dt_masks_read_masks_history() and attach it to dt_develop_t.history items,dt_dev_pop_history_items_ext(),dt_dev_history_gui_update(),dt_dev_pixelpipe_cache_peek_gui()dt_dev_pixelpipe_create_nodes(),dt_develop_t.history to dt_dev_pixelpipe_iop_t pipe pieces (nodes) through dt_dev_pixelpipe_synch_all() (update hashes internally),dt_iop_buffer_dsc_t, auto-disabling incompatible modules, through dt_dev_pixelpipe_propagate_formats(),_seal_opencl_cache_policy(),dt_dev_pixelpipe_t.backbuf (this is descriptive, there is no buffer copy).dt_dev_add_history_item() :dt_develop_t.forms,dt_dev_pixelpipe_t.backbuf,Notes: XMP files are interfaced with the library database, they are never used directly.
src/ is sorted so that a file's directory answers two questions without opening it: what layer is it and does it hold state. Both are enforced (see § The rules); neither is a convention you have to remember.
Layers run low to high. A file may include from its own layer or below, never above.
| layer | directory | holds state? | what belongs there |
|---|---|---|---|
| 0 | system/ | no — guaranteed | platform and machine facts: allocation, SIMD, OpenMP, thread wrappers, dynamic loading, resource limits, the arena allocator, device-scaled cairo surface arithmetic |
| 0 | win/, external/ | — | Windows shims for POSIX; vendored third-party |
| 1 | math/ | no — guaranteed | pure algorithms: splines, expression evaluation, topological sort |
| 1 | common/ | yes | application services: database, caches, config, logging, image metadata, styles, tags |
| 1 | colorprofiles/ | yes | ICC profiles: the LittleCMS2 work, the camera matrices, the transform engine, and the application-wide profile list |
| 2 | pixel/ | no file holds its own | pixel maths: filters, wavelets, interpolation, colour transforms |
| 3 | control/ | yes | jobs, signals, progress, the control loop |
| 4 | widgets/ | two named registries only | the GTK widget set — reusable, and may not include gui/ |
| 4 | gui/ | yes | this application's windows, panels, theme, screen metrics |
| 5 | develop/ | yes | history, pipeline, masks |
| 6 | iop/, imageio/ | yes | image operations; codecs and export |
| 7 | libs/, views/ | yes | panel modules and views |
| 9 | src/ root | yes | darktable.c, the orchestrator |
Measured with tools/statelessness_audit.py: 739 files, 316 headers, 0 include cycles, 217 layering violations. Re-measured 2026-09-29 with tools/include_graph.py --summary: 814 nodes, 361 headers, cycles 0, layering_violations 183 (and tools/include_baseline.txt agrees). Note statelessness_audit.py needs a populated build-debug/ — it reads the object files — and errors out with no .o files under .../build-debug otherwise, which is why the cheaper include_graph.py is the one to reach for when only the graph numbers are wanted.
A stateless module can be reused, tested, threaded and ported without dragging the application behind it. The point of sorting by it is that the check becomes mechanical: anything built only from system/ and math/ is stateless too, and nobody has to re-derive that per file.
That inference is only sound while those directories stay closed, which is why tools/check_module_boundaries.sh pins what system/ may include. One exception would cost every reader the check the arrangement exists to avoid.
**widgets/ cannot be stateless.** Ten of its files hold GObject type registration — G_DEFINE_TYPE's _private_offset/_parent_class/static_g_define_type_id, the cached g_type_register_static id, g_signal_new ids. Registering a type once per process and caching the id is what defining a GTK widget class is. The rule instead is: no state outside widget_settings.c and accelerators.c, its two named registries.
**common/ is where state is allowed to live.** It is not a dumping ground — it is the answer to "this needs a database connection / a config file / a cache". If something there holds no state, it belongs in system/, math/ or a module of its own.
Five gates run in CI. Each exists because the thing it checks broke something that no other gate could see, and every one of them fails the build rather than warns.
They do not all run on every build, which this said until it was measured (2026-09-29): check_conditional_includes.sh runs on pull requests only, and check_unused_includes.sh on pull requests only and only in the LLVM20 + skiptest cell — 1 of 13. Both gate a diff and a diff needs a base ref, so the restriction is right; the claim was not. A push to master is checked by three of the five. doc/ci.md has the table, the merge gate's design, and the holes the same audit found.
| gate | rule |
|---|---|
check_layering.sh | no include from a higher layer; no include cycles. A ratchet — the count may fall, never rise |
check_module_boundaries.sh | system/ includes nothing that could bring state with it; widgets/ does not include gui/ |
check_statelessness.sh | system/ and math/ hold no state; widgets/ holds none outside its two registries. Measured from the compiled objects |
check_unused_includes.sh | no newly added include is unused. Renames and pre-existing findings are reported, not gated |
check_conditional_includes.sh | no newly added include sits inside a conditional block |
Plus the two invariants that predate them: no #pragma once (explicit guards make cyclic includes greppable; see include-graph.md), and **darktable.h has a re-inclusion tripwire rather than a guard** — if you hit it, give the file the specific library it needs.
A header includes only what its own declarations need. Everything else belongs in the .c.
A header that includes more becomes a supply line its consumers never asked for and cannot see. They compile because something upstream happened to pull in what they use, and the day anyone tidies that include away, the breakage surfaces somewhere else entirely — in a file nobody touched. Two real instances:
gui/gtk.h broke a dozen IOPs, because it had been the only thing pulling sqlite3.h in ahead of common/points.h, whose vendored SFMT #define N then collided with sqlite3_compileoption_get(int N);widgets/label.h from widgets/dialog.c removed dt_free, which had been arriving through label.h → system/mem_alloc.h in a file that named neither.Implementations do not belong in headers either: five static inline helpers in widgets/label.h forced four extra includes on every one of its ~30 consumers. The one legitimate exception is a header whose published interface is inline code (widgets/draw.h), which necessarily includes what that code calls. Keep those rare and honest — they are a performance trade, not a convenience.
When a lower layer needs something only a higher layer knows, do not include upward. Declare a handler and let the higher layer register it:
Unregistered must be a defined no-op, so headless runs and early startup work without a single "is there a GUI?" test. widgets/widget_settings.h is the worked example — cursor shape, toast messages, natural width, per-widget persistence and the root window all arrive this way.
For a value rather than a callback, push it down instead: the application resolves it once and stores it where the consumer already lives. Screen DPI works like that — widget_settings owns it because every widget reads it on every draw, and gui/screen_metrics.c forwards under the dt_screen_* names for the few readers below layer 4.
| tool | answers |
|---|---|
include_graph.py --summary | cycles, layering violations, closure sizes |
statelessness_audit.py --dir src/x | which files hold or reach state, and the chain that gets there |
header_consumers.py <header> | what each includer actually takes from a header — its own symbols vs. what it merely forwards |
header_includes_audit.py <header> | includes a header does not need, and types it reaches transitively |
fix_missing_includes.py | reads a compiler log, adds the header declaring each missing symbol |
check_windows_syntax.sh | syntax-checks changed files under the Windows preprocessor via MinGW, without a Windows build |
A green build on one configuration proves very little here.
nofeatures (mirrors CI's reduced dependency set), Debug, and a clang Debug. Compilers disagree in ways that matter — the statelessness gate once passed three GCC builds and failed the LLVM job, because clang names function-local statics <function>.<variable> where GCC uses <variable>.<n>._WIN32, __APPLE__ or GDK_WINDOWING_* is invisible on a Linux desktop. check_windows_syntax.sh covers the preprocessor-branch case for about a second per file.tools/check_it_runs.sh exports one small PNG, exactly as CI's "Check if it runs" step does. Every other check here is static and none of them can see a double free or a use-after-free. A change that passed four build configurations and every gate has already aborted all eight CI runners with heap corruption.ctags finds before and after. Line counts and build status both miss silent deletion — a view truncated to its include block still compiles, it just stops being a view.Silent supply. Most breakage in this series was not the code that changed; it was code that had been compiling by accident and stopped. Expect a structural change to surface unrelated files, and read that as the point of the exercise rather than as collateral damage.
Gates that tax the wrong thing. A gate scoped to "every file this change touches" reports a file's entire inherited backlog when a refactor rewrites one include line in 180 files, and a gate that cries wolf gets switched off. Scope to what the change adds; report the rest as notes. The same applies to renames — a moved header is an added line, and demanding the file justify an include it has always carried is how mechanical work starts dragging cleanup along.
Roughly in dependency order. Layering violations (183 as of 2026-09-29, 217 when this was written) fall as these land.
Extract database, caches, metadata from common/. common/ is 63 translation units and remains the largest undifferentiated module. Database access in particular should be behind one API with its own per-image locking, so thread-safety stops depending on every caller remembering to lock the image cache first (see § Database).
Move the GUI half of imageio/. imageio/format/, imageio/storage/ and imageio_module.c are export dialogs and settings widgets, not codecs. They belong in gui/.
Split develop/imageop_math.h. The curve-fitting maths is pure and belongs in math/.
Remove dev->proxy. It is an anti-pattern that breaks modularity by design — a grab-bag of function pointers letting anything reach anything.
Make pixel/ stateless. No file there holds state of its own; all 13 that reach state do so through common/opencl.c (the device registry) or caches/pixelpipe_cache.c. Inverting those two dependencies would make the tree's whole pixel-maths layer stateless.
Close system/ fully. It no longer includes anything outside itself, but the gate is what guarantees that rather than the structure. Nothing to do unless it regresses.
These concern the runtime architecture rather than the module layout above: known contention, and directions the pipeline could take. They are design intent and measured problems, not descriptions of finished work — anything marked open is unclaimed.
dt_dev_history_item_t is now refcounted the same way dt_masks_form_t already was (see masks_history.c): dt_dev_history_item_create() is the only constructor (initializes refcount to 1), dt_dev_history_item_ref() takes a reference, dt_dev_free_history_item() is unref-and-free-at-zero, and dt_dev_history_cow_touch() clones a shared item before mutating it in place (mirrors dt_masks_cow_touch()). This exists so that any future code that needs to hold a snapshot of dev->history across a slow operation (pipe resync, DB write job) can do so without a deep copy and without racing a GUI-thread mutation of the same item.
The refcounting was introduced to attack a real, measured problem: dt_dev_pixelpipe_change() (the worker-thread resync function, called from dt_dev_darkroom_pipeline()) holds dev->history_mutex as a reader for the entire O(nodes × history) resync of a pipe, which routinely takes tens of ms and, under mask-heavy history stacks combined with active editing, has been measured over 200ms. Because dt_dev_add_history_item_ext() (the GUI-thread commit path) needs the writer lock on the same history_mutex, and the writer-preferring rwlock policy blocks new readers once a writer is queued, every GUI edit (a scroll on exposure, a mask drag) can stall for the full duration of whatever resync the worker thread happens to be mid-flight on. This is directly observable with the named-rwlock diagnostic added to dt_pthread_rwlock_t (system/dtpthread.h: dt_pthread_rwlock_set_name() + wait-time logging, opt-in per lock — dev->history_mutex is named in dt_dev_init()) combined with -d history.
Status: fixed. dt_dev_pixelpipe_change() now resyncs against a dt_dev_history_snapshot_t (dt_dev_history_snapshot_take() / _release(), dev_history.h): the read lock is held only to copy the list cells, take one reference per item and capture history_end and the history hash — the three fields dt_dev_set_history_end_ext() writes together — and is released before any commit_params() runs. The writer's copy-on-write gate above is what makes that sound: an item a snapshot holds has refcount > 1, so dt_dev_history_cow_touch() clones it rather than mutating it under the reader.
Measured on the load-time resync of _DSC9410.NEF (47 history items, mask group with a 57-node brush), -d history, same machine, same scratch config, before and after:
| lock held (preview pipe) | lock held (full pipe) | resync work | |
|---|---|---|---|
| before | 204.66 ms | 226.75 ms | (inside the hold) |
| after | < 1 ms (below the print threshold) | < 1 ms | 174.63 ms / 158.52 ms, lock-free |
The compute cost is unchanged — the resync is as expensive as it was — but nothing waits on it any more. pipe->last_history_item became an atomic pointer the pipe holds a reference on, exchanged with dt_atomic_exch_ptr() (added to system/atomic.h), so the marker that synch_top bounds its work with stays valid across a copy-on-write clone landing from the GUI thread while the worker writes the slot outside the lock. tests/unittests/test_history_snapshot.c pins the snapshot/COW/refcount contract. See CLAUDE.md, "History items are refcounted; the pipe
resyncs against a snapshot, not under `history_mutex`", for the three design constraints that are not obvious from the code.
The second long reader of this lock, the async DB write job (_dt_dev_write_history_job_run), took the same cure: _history_write_state_take() freezes the snapshot plus a deep copy of dev->iop_order_list under a brief read lock, and _write_history_from_state() performs the full history+masks rewrite lock-free. Its one extra constraint is that the coalescing flag history_write_pending is cleared at snapshot time, under the lock, so a commit landing during the rewrite queues its own write instead of being silently dropped.
The following are proposals, enabled by the pipeline-cache design described above but not implemented. They are recorded because the cache rewrite was done partly to make them possible.
The history is so far incrusted into the dt_develop_t object, with its own history_end and history_mutex lock. It mixes both module parameters history, and masks/forms history in a weird fashion. The history use to be scattered all over the software, with parts handled in SQL and parts handled in C. Now that everything is handled in C and the history handling methods have been contained in history.c and dev_history.c, it might be a good idea to move it entirely out of the dt_develop_t object to have it managed globally, like the pipeline cache, but behind an API that allows to track precisely the lifecycle of data and concurrent accesses. Writing history to database and to XMP is currently handled in separate, short-lived threads, this completely enclosed architecture would allow to have all history tasks (including reading) handled in parallel, while ensuring thread safety.
The current architecture has the pipeline rendering triggered implicitely from history changes. This means we need to resync the history for all relevant modules all the time, and relevant modules, when working with masks, are all modules. This puts a solid 30-50 ms delay before starting any pipeline on history changes, or figuring out that no recompute is needed. But the current architecture of the pipeline cache allows us to start and stop pipelines absolutely how we want, because we can probe and fetch the cachelines from any thread. It means that, when changing GUI parameters in modules, modules could self-start their process()/process_cl() directly in pipeline, calling dt_iop_commit_params() on themselves, and fetching their own input straight from cache, therefore bypassing history (history would still have to be written, but the pipeline render wouldn't have to wait for it in its own thread). drawlayer.c would greatly benefit from such low latency.
A direct module -> pipeline trigger would also avoid the messy business of having to handle mask preview GUI states within the pipeline recursion, which means that we have to hack the piece->global_hash to account for GUI states and properly recompute pipelines when switching on/off mask previews. That introduces lots of edge-cases to handle through heuristics and bypasses. Raster-mask providers already publish dedicated side-band cachelines, independently from their image outputs; the remaining architectural step would be a specialized pipeline that only runs the distort_mask() methods instead of performing those transforms while the consuming module blends. Note that color-pickers have also been completely removed from the rendering pixel pipeline and now deal directly with cachelines, from the GUI thread, which avoids having to recompute a pipe just to refresh their values.
The new architecture doesn't force modules to take their input from the previous module output either, now they can take input from any module output in the pipeline as long as we know the global hash of the module, to fetch its output buffer on the cache. So we could have a new module, masking & merging that could take input from several modules and blend them over each other with alpha, meaning we could have parallel branches within pipelines and not stay limited by single sequences of modules. Along with the new nodal graph viewer, that would allow a full nodal workflow.
Found b69f8864d4, 2026-06-25. Verified against 42eca0e8fe, 2026-09-29.
Ansel inherited the Darktable practice of entangling every application layer (GUI, pipeline, history, database) and importing the whole software into the whole software through #include "darktable.h". That voided the modularity principle, caused bugs and data races, and made maintenance prone to edge effects in an application that is heavily asynchronous and parallel.
That specific problem is now largely closed, and this paragraph used to say otherwise. Measured at
42eca0e8fe(2026-09-29): 13 files in all ofsrc/includedarktable.h, out of 816 — fivemain()s and eight top-level translation units, which is exactly the shape the goal below prescribes. No header includes it at all, and the file itself is down to ten includes. Seven headers now carry comments actively routing callers away from it. Written in the present tense, this section sent a reader to go and fix something the darktable.h strip had already done. The risk today is regression, not remediation — keep it that way.
The Ansel codebase should move toward more enclosed modularity, making data structure private to each translation unit and exposing only API to the outside (getters/setters/init/cleanup). Direct value changes on data not owned by the current TU are forbidden. The dependency graph should be simplified and only a minimal set of #include should be kept per TU. In particular, src/darktable.h inherits from lower-level modules and lower-level modules must not inherit it: it has stopped being the glue of all common helpers, and nothing should make it that again.
CRUD operations should have one central entry point for the whole software and run only once, for as long as user didn't send new input, so the data lifecycle is legible and cacheable.
Since every data flow in the software is a pipeline, issues should be tracked to their root cause by climbing the call tree up until the source is found, instead of being fixed where they are visible.