Command line
lectes comes with a lectes command that scans a file, or stdin, with a grammar
file and prints the tokens. It's handy for developing a grammar and for feeding
tokens to other tools.
Running lectes with no arguments prints the help, as lectes --help does.
See Installation for how to put lectes on your PATH.
Grammar files
A grammar file has one rule per line: a name, whitespace, then the regex, which
is the rest of the line. Blank lines and lines starting with # are ignored.
Rules are tried in order, and the longest match wins, so the order only
matters when two rules match the same text.
By convention, grammar files use the .lectes extension.
Try it in the playground: paste a grammar into lectes.dev to see its tokens as you type, and why each rule won.
Rule names start with a letter or underscore and contain only letters, digits
and underscores. Regexes use Python's re
syntax and are written as-is, with no quoting or escaping beyond what the regex
itself needs.
The same format can be loaded from Python with Configuration.from_text; see
Usage.
Output
By default each token is printed on its own line, as its line:column, its
rule name and its literal:
$ echo '12 + x' | lectes arithmetic.lectes
1:1 INT '12'
1:3 WS ' '
1:4 PLUS '+'
1:5 WS ' '
1:6 ID 'x'
1:7 WS '\n'
Literals are shown with Python's repr, so whitespace and other special
characters are visible.
With --format json, each token is printed as a JSON object on its own line
(JSON Lines), ready for jq or another program:
$ echo '12 + x' | lectes arithmetic.lectes --format json
{"name": "INT", "literal": "12", "offset": 0, "line": 1, "column": 1}
{"name": "WS", "literal": " ", "offset": 2, "line": 1, "column": 3}
...
offset is 0-based and counts characters; line and column are 1-based.
Unmatched text
By default, scanning stops at the first text that no rule matches. The tokens before it have already been printed, and the unmatched text is reported on stderr:
With --ignore-unmatched, every run of unmatched text is reported on stderr
and scanning continues. Unmatched text is always reported on stderr, in both
output formats, so stdout only ever holds tokens.
Grammar errors and warnings
Problems in the grammar file are reported on stderr with their line number, in a form most editors can jump to. Every error is reported, not just the first one:
$ lectes broken.lectes input.txt
broken.lectes:1: error: invalid rule name '9X'
broken.lectes:3: error: rule 'ID' has no pattern
Warnings point out rules that are valid but probably not what you meant, such as a rule that can match the empty string. They're reported the same way, but scanning goes ahead.
Use --check to validate a grammar without scanning anything.
Options
| Option | Description |
|---|---|
INPUT |
File to scan. Reads stdin when omitted or -. |
--format {text,json} |
Output format; text by default. |
--ignore-unmatched |
Report unmatched text on stderr and keep scanning. |
--encoding ENCODING |
Encoding of the input; utf-8 by default. Grammar files are always UTF-8. |
--check |
Only validate the grammar. |
--debug |
Print the scanner's debug events to stderr. |
--version |
Print the version and exit. |
Exit codes
| Code | Meaning |
|---|---|
0 |
Success. With --ignore-unmatched, unmatched text doesn't change this. |
1 |
The input has text that no rule matches. |
2 |
A usage error, an invalid grammar, or a file that can't be read or decoded. |