C and C++
Galley-generated parsers can be consumed from C and C++ through a small application-binary interface: Galley compiles a generated parser into a shared library (lib<name>.dylib / .so) together with the C header bindings/c/galley.h. Complete, runnable consumers live in examples/c and examples/cpp; both are built and executed by CI on every push.
Procedures
Grammars can use @hook_name annotations on RHS occurrences. When Galley's --emit-metadata flag is passed during generation, it produces a procedures.zig alongside the parser with extern declarations for every hook. When procedures.c (or procedures.cpp for C++) lives next to the parser, the consumer build compiles it automatically; otherwise pass its location explicitly with -Dprocedures-c-source=procedures.c. Reduction hooks keep their reduction_<VariableName> names (plus the general reduction); author-defined grammar hooks are declared as hook_<name>, namespacing them away from unrelated symbols:
/* procedures.c */
#include <galley.h>
#include <stdio.h>
void reduction_Pair(void *args) {
GalleySession *session = galley_procedure_session(args);
GalleyNodeAddress node = galley_procedure_current_node(args);
const char *text = NULL;
size_t len = 0;
unsigned line = 0, column = 0;
if (session == NULL) return;
galley_node_text(session, node, &text, &len);
galley_node_line_column(session, node, &line, &column);
fprintf(stderr, "Pair %.*s (%u children) at %u:%u\n",
(int)len, text, galley_node_child_count(session, node), line, column);
}
void reduction_KeyTail(void *args) {
galley_procedure_drop_if_empty(args);
}
void hook_print(void *args) {
GalleySession *session = galley_procedure_session(args);
GalleyNodeAddress node = galley_procedure_current_node(args);
const char *text = NULL;
size_t len = 0;
unsigned line = 0, column = 0;
if (session == NULL) return;
galley_node_text(session, node, &text, &len);
galley_node_line_column(session, node, &line, &column);
fprintf(stderr, "@print \"%.*s\" at %u:%u\n", (int)len, text, line, column);
}Each hook receives an opaque ProcedureArguments pointer. Recover the parsing session with galley_procedure_session and inspect nodes with the ordinary galley_node_* functions. Drop/replace the current node with galley_procedure_drop_* / galley_procedure_replace_with_children; those talk to the parser through args.node_address and are not the same as galley_tree_remove_self.
Semantic payloads remain unavailable through the C API.
Error Messages
Messages are customizable without any Zig: message overrides (fixed strings with placeholders) and host-side renderers (dynamic callbacks) cover virtually every need — see the two sections below.
The advanced escape hatch remains a generated Zig hook file: run galley --fill-error-messages examples/c, edit any hook body (for example syntax_error_ll_Number__expected_generative_terminal_digit); when ll_error_messages.zig lives next to the parser the consumer build picks it up automatically, otherwise pass it explicitly:
"-Derror-messages-zig-source=/path/to/ll_error_messages.zig"galley_diagnostic_message then returns the text your hooks render; without the flag (or for un-customized grammars) it returns the built-in generic renderer output. The _ansi accessor always renders generically. LR grammars use the same flow with lr_error_messages.zig and syntax_error_lr_* hook names.
Message Overrides
To replace messages with fixed strings — no Zig file at all — register overrides keyed by structured identity: the innermost in-progress variable name (for example "Number"), or "*" for every syntax and indentation error. Variable keys win over "*"; overrides take priority over hooks. Placeholders expand against the failing diagnostic:
galley_session_set_message_override(session,
"Number", sizeof("Number") - 1,
"expected a number after ':' (digits only) at line {line}",
sizeof("expected a number after ':' (digits only) at line {line}") - 1);Both strings are copied; overrides persist for the session's lifetime.
Placeholders inside override messages expand against the failing diagnostic: {line}, {column}, {unexpected}, {expected} (rendered as 'a', 'b'), and {context} (innermost-first chain joined with <~). Unknown names pass through untouched.
Build Model
Consumers drive two commands from whatever build system they prefer — no Galley-side build knowledge is required:
Generate the parser from a grammar with the generator CLI (operating on a language directory containing
ll.grmand/orlr.grm; boilerplate modules are created automatically):sh<galley>/zig-out/bin/galley --parser-type ll /path/to/language-dir # → /path/to/language-dir/_ll-parser.zig (--parser-type lr → _lr-parser.zig)Compile the generated parser into a shared library with Galley's generic consumer build file, directly next to the grammar:
shzig build --build-file <galley>/bindings/c/consumer/build.zig \ "-Dlanguage-dir=/path/to/language-dir" \ "-Dlib-name=mylang" \ "-Doutput=libmylang.so" \ "-Doptimize=ReleaseFast" \ --prefix /path/to/language-dir install # → /path/to/language-dir/libmylang.so (no lib/ layer, no header; # read galley.h from <galley>/bindings/c, or pass -Dinstall-header)Both parser families work identically through this ABI: the consumer locates
_ll-parser.zigvs_lr-parser.zigin the language dir and infers the family from the filename (-Dparser-sourceplus-Dparser-typeonly for non-standard filenames and layouts). One library embeds one parser.
Generation-time options come from config.zig in the language directory; CLI flags edit its constants in place.
The consumer build infers every language-owned source next to the parser when no explicit flag is given — config.zig, procedures.zig, {ll,lr}_error_messages.zig (when present, otherwise the built-in template), and a procedures.c or procedures.cpp implementation when present. Explicit flags (-Dconfig-zig-source, -Dprocedures-zig-source, -Dprocedures-c-source / -Dprocedures-object, -Derror-messages-zig-source, -Dparser-source / -Dparser-type) override inference and exist only for non-standard layouts where those files live elsewhere. The reference examples/c and examples/cpp builds pass only parser-source and rely on inference for the rest.
What the examples' CMake does
Both examples wire steps 1–2 into CMake so a plain cmake -S examples/c -B build -DCMAKE_BUILD_TYPE=Release -DGALLEY_CHECKOUT="$PWD" && cmake --build build builds its CLI, generates the parser from the example's own ll.grm, compiles the library next to the grammar, builds build/bin/demo and build/bin/benchmark, and runs nothing else. Generation also re-runs automatically whenever ll.grm or config.zig changes. Without GALLEY_CHECKOUT the configure step fails loudly (for convenience, GALLEY_CHECKOUT=$(examples/scripts/fetch-galley.sh) fetches one).
Useful variables:
| Variable | Purpose |
|---|---|
GALLEY_CHECKOUT | Existing Galley working tree (required) |
Generated files (_ll-parser.zig, config.zig, procedures.zig) and the grammar library (libkeyvalue-c.*, libbenchmark-c.*) live in the example directory and are gitignored. After a build, build/bin/ contains demo and benchmark — the Galley CLI stays inside its own tree.
Both example directories also emit compile_commands.json next to their sources and ship a .clangd fallback, so editors resolve <galley.h> and offer completion before the first build.
Passing a file path as the only argument parses that file and nothing else (exit status reports success; failures print a diagnostic):
./build/bin/demo path/to/input.fileRuntime Concepts
Sessions
GalleySession *session = galley_session_create();
/* or with options: */
const GalleyCOptions options = { .max_errors = 10 };
GalleySession *session = galley_session_create_ex(&options);
/* ... */
galley_session_destroy(session);Sessions are not thread-safe — use one per thread or guard externally. All result data (node addresses, text pointers, diagnostic strings) remains valid until the next parse on the same session or session destruction.
Parsing
long long parsed = galley_parse_sentinel(session, input); /* NUL-terminated */
long long parsed = galley_parse(session, data, len); /* arbitrary bytes */
long long parsed = galley_parse_file(session, "file.json"); /* from disk */Returns the number of bytes parsed on success, or a negative galley_error_* code (galley_status_string renders any code).
Walking the AST
Node handles are stable byte indices (GalleyNodeAddress); editing never invalidates them. Walk depth-first through the shared walker rather than hand-rolling recursion, so order and depths match every binding:
GalleyWalker *walker = galley_walker_create(session, galley_root_node(session), 0);
GalleyNodeAddress n; unsigned int depth; int is_semantic_error;
while (galley_walker_next(walker, &n, &depth, &is_semantic_error)) {
const char *name_data; size_t name_len;
const char *text_data; size_t text_len;
unsigned int line = 0, column = 0;
galley_node_symbol_name(session, n, &name_data, &name_len);
galley_node_text(session, n, &text_data, &text_len);
galley_node_line_column(session, n, &line, &column);
}
galley_walker_destroy(walker);galley_walker_skip_children prunes the last yielded node's children, and a nonzero third galley_walker_create argument prunes subtrees rooted at semantic-error nodes. Destroy the walker before destroying the session or parsing again.
galley_node_first_child, galley_node_next_sibling, galley_node_child_count, galley_node_last_child, galley_node_prior_sibling, galley_node_parent, galley_node_span, and galley_node_variable_index complete the read surface. galley_tree_snapshot reads the same columns for every node in one call into caller-owned flat arrays (parent, first child, next sibling, child count, variable index, span start, span length); it returns the node count and writes up to capacity entries, so size with galley_node_count first and pass null for columns you do not need. Spans index the retained input, readable in one call with galley_last_input.
Editing the Tree
Chains passed to edit functions must be detached orphans; edits never invalidate other addresses.
GalleyNodeAddress head;
galley_tree_clean_children(session, parent, &head);
galley_tree_append_children(session, parent, head);
galley_tree_insert_before(session, target, chain);
galley_tree_insert_after(session, target, chain);
galley_tree_insert_children_at(session, parent, index, chain);
galley_tree_remove_siblings(session, node, count, &head);
galley_tree_remove_self(session, node, &head);
galley_tree_remove_children_at(session, parent, index, count, &head);
galley_tree_promote_children_over_wrapper(session, wrapper, &head);
galley_tree_unlink_wrapper(session, wrapper);Diagnostics
When a parse fails, structured information is available until the next parse:
if (galley_has_diagnostic(session)) {
long long kind = galley_diagnostic_kind(session); /* none/syntax/indentation/semantic */
unsigned int line, column;
galley_diagnostic_position(session, &line, &column);
const char *msg;
galley_diagnostic_message(session, &msg); /* plain text */
galley_diagnostic_message_ansi(session, &msg); /* colored */
long long count = galley_diagnostic_expected_count(session);
for (long long i = 0; i < count; ++i) {
const char *tok; size_t len;
galley_diagnostic_expected_at(session, i, &tok, &len);
}
/* context chain: galley_diagnostic_context_count/_at */
/* unexpected token: galley_diagnostic_unexpected_token */
/* indentation details: galley_diagnostic_indentation */
/* recovery target: galley_diagnostic_recovery_* */
}galley_syntax_error_count reports how many errors a recovery-enabled parse recorded.
Recovery-enabled parses retain every diagnostic they record, addressable by index (0-based, in recording order) until the next parse:
long long recorded = galley_recorded_diagnostic_count(session);
for (long long i = 0; i < recorded; ++i) {
unsigned int line, column;
galley_recorded_diagnostic_position(session, i, &line, &column);
long long kind = galley_recorded_diagnostic_kind(session, i);
/* plus recorded_{unexpected_token,expected_count,expected_token,
context_count,context_name,indentation,recovery_*}, mirroring the
singular accessors above; messages render generically via
galley_recorded_diagnostic_message */
}The singular accessors report the most recent diagnostic.
Parser Metadata
galley_parser_type, galley_error_recovery_mode, galley_has_ast, galley_has_procedures, galley_source_retention_enabled, galley_has_position_tracking, galley_has_input_streaming, galley_uses_verbatim, galley_stack_overflow_recovery_available, and the grammar symbol table (galley_symbol_count / _name / _is_terminal, galley_variable_count / _name) let embedders introspect exactly what the library was built with.
Storage Notes
AST nodes live in non-relocating storage (a reserved contiguous region on macOS/Linux/BSD, fixed segments elsewhere), which is why node addresses and pointers derived from them are stable across allocations. Node storage can be preallocated with galley_reserve_nodes; galley_node_capacity reports the current capacity.
Development builds
Every green CI run uploads per-platform kits (galley-c-<version>-<platform>.tar.gz) as workflow artifacts (Actions → the run → Artifacts → pkg-c), holding the generator binary, galley.h, and the compile inputs — no repo checkout needed. Each kit ships a README with the two commands; you still need a Zig 0.16.0+ toolchain and a C compiler:
mkdir -p galley-c && tar xzf galley-c-<version>-linux-x64.tar.gz -C galley-c --strip-components=1
./galley-c/bin/galley --emit-metadata <language-dir>
zig build --build-file galley-c/share/galley/compile-kit/build.zig \
-Dlanguage-dir=<language-dir> -Dlib-name=<name> \
-Doutput=lib<name>.so -Doptimize=ReleaseFast \
--prefix <language-dir> installVersioned releases carry the same kits under versioned names for anything durable. One kit serves C and C++ alike.
Related Pages
- Using Galley as a Library — the language-directory generation flow in detail
- Rust and Go — bindings over the same shared library
- Grammar Guidelines
- Architecture