Skip to content

Usage

Defining the scanning rules

You have to define a set of Rules based on which the scanner will operate.
A Rule is just a named regex pattern.

from lectes import Rule, Regex, Configuration

config = Configuration(
  [
    Rule(name="FOR", regex=Regex("for")),
    Rule(name="IN", regex=Regex("in")),
    Rule(name="ID", regex=Regex("[a-zA-Z_][a-zA-Z_0-9]*")),
    Rule(name="COLON", regex=Regex(":")),
    Rule(name="WHITESPACE", regex=Regex("( )")),
  ]
)

Defining the rules in a grammar

The same rules can be written as grammar text, one rule per line: a name, whitespace, then the regex, which is the rest of the line. Blank lines and lines starting with # are ignored. This is the format the lectes command reads.

from lectes import Configuration

config = Configuration.from_text(
  r"""
  # a tiny language
  FOR         for
  IN          in
  ID          [a-zA-Z_][a-zA-Z_0-9]*
  COLON       :
  WHITESPACE  ( )
  """
)

Use a raw string, or read the grammar from a file, so that backslashes in the regexes reach the parser unchanged.

Try it in the playground: lectes.dev scans as you type, so it's a quick way to work a grammar out before putting it in your code. Its Export as Python button gives you the Configuration.from_text call.

If the text has errors, from_text raises a GrammarError. Its problems list holds every error found, as (line, message) pairs, so all of them can be fixed at once. line is None for problems that aren't tied to a line, such as an empty grammar.

from lectes import GrammarError

try:
  Configuration.from_text("9X a\nID")
except GrammarError as e:
  print(e.problems)  # [(1, "invalid rule name '9X'"), (2, "rule 'ID' has no pattern")]

Rules that are valid but probably not what was intended, such as one that can match the empty string, emit a GrammarWarning through the warnings module. Each warning carries its line and message.

import warnings
from lectes import GrammarWarning

with warnings.catch_warnings(record=True) as caught:
  warnings.simplefilter("always", GrammarWarning)
  config = Configuration.from_text("OPT x?")

for warning in caught:
  print(warning.message.line, warning.message.message)

Scanning

The scanner only requires a Configuration of Rules to be initialized, it can then scan any given text based on that configuration.

from lectes import Rule, Regex, Configuration, Scanner

config = Configuration(
  [
    Rule(name="FOR", regex=Regex("for")),
    Rule(name="IN", regex=Regex("in")),
    Rule(name="ID", regex=Regex("[a-zA-Z_][a-zA-Z_0-9]*")),
    Rule(name="COLON", regex=Regex(":")),
    Rule(name="WHITESPACE", regex=Regex("( )")),
  ]
)

scanner = Scanner(config)

for token in scanner.scan("for var in array:"):
  print(token)

Defining custom handlers

When a rule is matched, the default behaviour of the scanner is to yield a Token object. A token holds the matched rule, the matched literal and the location where the literal starts in the text (its offset, line and column).

This behaviour can be tweaked by defining custom handlers for individual rules. A handler receives the matched Token; whatever it returns is yielded by the scanner, unless it returns None, in which case the token is skipped.

from lectes import Rule, Regex, Configuration, Scanner, Token

config = Configuration(
  [
    Rule(name="FOR", regex=Regex("for")),
    Rule(name="IN", regex=Regex("in")),
    Rule(name="ID", regex=Regex("[a-zA-Z_][a-zA-Z_0-9]*")),
    Rule(name="COLON", regex=Regex(":")),
    Rule(name="WHITESPACE", regex=Regex("( )")),
  ]
)

def whitespace_handler(token: Token) -> None:
  return

ids = []

def id_handler(token: Token) -> Token:
  ids.append(token.literal)
  return token

# You don't have to return a Token
def for_handler(token: Token) -> dict:
  return {"matched": token.literal, "line": token.location.line}

scanner = Scanner(config)
scanner.set_handler(config.rules[0], for_handler)
scanner.set_handler(config.rules[2], id_handler)
scanner.set_handler(config.rules[4], whitespace_handler)

for token in scanner.scan("for var in array:"):
  print(token)

print(ids)

Handling unmatched text

By default, the scanner raises an UnmatchedTextError when part of the text does not match any rule. Each contiguous run of unmatched text is reported once, together with the location where it starts. Since scan is a generator, the tokens before the unmatched text have already been yielded when the error is raised.

from lectes import UnmatchedTextError

try:
  tokens = list(scanner.scan("for var in array?"))
except UnmatchedTextError as e:
  print(e)  # unmatched text '?' at line 1, column 17
  print(e.unmatched.text, e.unmatched.location.line, e.unmatched.location.column)

To skip unmatched text instead, create the scanner with ignore_unmatched=True.

scanner = Scanner(config, ignore_unmatched=True)

To handle unmatched text yourself, define a custom handler. It receives an UnmatchedText with the text and its location, and replaces the default behaviour of raising an error. A custom handler cannot be combined with ignore_unmatched=True; trying to set one raises a ScannerConfigurationError.

from lectes import UnmatchedText

unmatched_text = []

def handler(unmatched: UnmatchedText) -> None:
  unmatched_text.append(unmatched.text)

scanner.set_unmatched_handler(handler)

Debugging

The debug argument can be passed in order to print the scanner's matches and unmatched text to stderr while scanning.

scanner = Scanner(config, debug=True)
DEBUG: rule ID matched: 'somevar'
DEBUG: unmatched: '@'

Only this scanner's events are printed; other scanners are not affected.

If the scanner has already been initialized without the debug flag, debug output can also be turned on by accessing the scanner's logger.

from lectes import LogLevel

scanner.logger().set_level(LogLevel.DEBUG)

Using standard logging

lectes logs through the standard logging module, under the lectes logger. To send the events of every scanner to your application's logging setup, configure that logger at DEBUG level instead of passing debug=True.

import logging

logging.basicConfig(format="%(name)s: %(message)s")
logging.getLogger("lectes").setLevel(logging.DEBUG)