Show HN: Yantra – an LALR(1) parser generator for C++

Hacker News by 5 min read 512x views
Show HN: Yantra – an LALR(1) parser generator for C++

Share Post

CI  MIT Version C++

Yantra Logo

Yantra is a mighty compiler compiler and LALR(1) parser generator written in C++, alongside the following center features:

  • An unified lexer
  • Built-in assistance for UNICODE/UTF8 input
  • Built-in AST builder
  • Built-in AST walker(s)
  • Bottom-up parsing (being LALR), and top-down strolling (traversal)
  • Multi-mode lexer, helpful for implementing nested multi-line comments, etc.
  • Lexer-driven (push-based) parser: says input one character at a period and feeds tokens to the parser as they complete, helpful for handling input as it arrives (e.g. from a socket).
  • An optional amalgamated mode, anywhere the complete parser is generated as a sole cpp file, alongside alongside a full-featured main() function.
  • Or, in non-amalgamated mode, the parser is generated as distinct .hpp and .cpp files, prepared to autumn into an existing project.

The name Yantra is Sanskrit for machine, as in state machine in this context.

Yantra has no requirements beyond the C++ norm library, so construction it is a plain CMake build:

git copy [email protected]:TantrixAuto/yantra.git cd yantra mkdir build && cd build cmake .. cmake --build .

This produces the ycc executable in bin/. Save a grammar file, hello.y:

start := stmts; stmts := stmts stmt; stmts := stmt; stmt := ID; ID := "[A-Za-z]+"; WS := "\s"!; 

Then create a parser from it:

bin/ycc -c ascii -f hello.y -a

This writes hello.cpp (an amalgamated, self-contained parser alongside its own main()) and hello.log. Compile it alongside any C++23 compiler:

# clang clang++ --std=c++23 -o hello hello.cpp # gcc g++ --std=c++23 -o hello hello.cpp # MSVC (cl.exe, from a Developer Command Prompt) cl /std:c++23 /EHsc /nologo hello.cpp

The grammar complete recognizes one or additional whitespace-separated alphabetic words. -s <string> feeds that cord immediately to the parser as input (as opposed to -f <filename>, which says from a file, or -i, which says interactively from the console):

# succeeds silently $ ./hello -s "hello world" $ echo $? 0
# -t1 prints the parsed AST $ ./hello -s "hello world" -t1 0:start_1(1:stmts_1(2:stmts_2(3:stmt_1(4:ID(hello))) 2:stmt_1(3:ID(world))) 1:_tEND())
# fails: ID lone matches letters, "123" isn't valid input for this grammar $ ./hello -s "hello 123" s1-err:?a1.in(001,007):TOKEN_ERROR{{token: }} hello 123 $ echo $? 1

Yantra parses the complete input into an AST first, then walks it top-down calling your semantic actions, dissimilar most parser generators, anywhere actions run bottom-up as all regulation is reduced. That ordering is what lets a genitor rule's act run before its children are visited. Save this as calc.y:

%class Calculator; start := expr; expr := expr(a) PLUS expr(b) %{ std::cout << "Adding" << std::endl; %} expr := NUMBER(N) %{ std::cout << "Number: " << N.text << std::endl; %} NUMBER := "\d+"; PLUS := "\+"; WS := "\s+"!; 

Generate and compile it the identical way as complete (bin/ycc -c ascii -f calc.y -a, afterward any of the three compiler commands), afterward run it:

$ ./calc -s "1 + 2 + 3" Adding Number: 1 Adding Number: 2 Number: 3

1 + 2 + 3 parses left-associatively as (1 + 2) + 3. So the external Adding, the base of the tree, prints first, followed by its remaining kid (Number: 1) and afterward its correct child, which is itself another Adding node alongside its own two children. A hand-written recursive-descent or bottom-up parser would have to build additional AST classes and a distinct strolling continue to get this ordering. Here it falls out of the grammar directly.

See the Build Instructions and Tutorial below for a genuine walk-through of the grammar syntax.

  • vs. Bison / Yacc / Lemon (the traditional LALR(1) family):
    • Lemon, from SQLite, is Yantra's straightforward stated inspiration.
    • These run semantic actions during parsing, as all regulation reduces, bottom-up.
    • Yantra continually builds the complete AST first, afterward walks it top-down in a distinct pass, so a genitor rule's act can run before its children are visited.
    • A sole grammar can additionally define additional than one walker (e.g. one that emits C++, another that emits Java, from the identical parse).
    • Getting either of those out of the Bison family method hand-building your own AST and walker on top.
  • vs. ANTLR:
    • ANTLR walks a fully-built parse tree too, but that comes for liberated from its LL(*) algorithm, which already builds the tree top-down as it parses.
    • Yantra gets the identical top-down stroll out of LALR(1), a bottom-up algorithm alongside no natural "whole tree exists yet" instant during parsing, during keeping LALR(1)'s period and area effectiveness complete adaptive LL(*).
    • Beyond that, Yantra targets C++ lone (ANTLR generates for many languages) and ships its own unified lexer alongside mode-stack assistance alternatively of a distinct lexer generator.
    • ANTLR's own generator tool is Java, so using it from a C++ project method adding a JVM to the build toolchain fair to run the generator. Yantra is a native C++ executable alongside no specified dependency.
    • ANTLR is far additional mature and extensively used. Yantra is a much smaller, newer, single-maintainer project.
  • vs. tree-sitter:
    • A distinct issue entirely. It's built for incremental, error-tolerant parsing embedded in editors and IDEs (what GitHub, Neovim, etc. use it for), not for generating a compiler/codegen backend.
    • Yantra doesn't do incremental reparsing and isn't trying to.

See Known Limitations for an honest catalog of what Yantra doesn't do yet.

The following are a set of key links to get acquainted alongside Yantra.

It is recommended that they be peruse in the stated order.

See https://github.com/TantrixAuto/lingo for standalone example project that uses yantra.

Language Server Extension

This is a tongue server expansion created by Raj Chaudhuri that provides syntax highlighting for Yantra records in vscode, qtcreator, and any another IDE that supports the Language Server Protocol.

https://github.com/rajware/yantra-language-server

Yantra is licensed under the MIT License.

Renji Panicker (@renjipanicker)

Other Article Hacker News
↑
Close Right Ads
Close Left Ads