scanner
Scanner
Scans a given text and returns tokens based on the provided configuration.
At each position the scanner picks the rule with the longest match; if
several rules match the same length, the one configured first wins. Within
a single rule, matching follows Python's re semantics (leftmost-first),
so a pattern like a|ab matches only a.
Text that no rule matches raises UnmatchedTextError by default. Pass
ignore_unmatched=True to skip it instead, or set a custom handler with
set_unmatched_handler.
Example
from lectes import Rule, Configuration, Regex, Scanner
config = Configuration(
[
Rule(name="FOR", regex=Regex("for")),
Rule(name="INT", regex=Regex("[1-9]+")),
Rule(name="ID", regex=Regex("[a-zA-Z][a-zA-Z0-9]*")),
Rule(name="WHITESPACE", regex=Regex("( )")),
]
)
scanner = Scanner(config)
program = "somevar in othervar for 9 let"
for token in scanner.scan(program):
print(token)
Source code in src/lectes/scanner/scanner.py
9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 | |
logger()
scan(text)
Scan the given text and yield tokens as they are recognized.
Each contiguous run of text not matched by any rule is passed to the
unmatched handler as a single UnmatchedText, once the run ends.
Raises:
| Type | Description |
|---|---|
UnmatchedTextError
|
if part of the text matches no rule, unless
|
Source code in src/lectes/scanner/scanner.py
set_handler(rule, handler)
Set the given function as the handler that executes when a string is matched against rule.
The handler receives the matched Token. Whatever it returns, except None,
is yielded by scan; returning None skips the token.
Source code in src/lectes/scanner/scanner.py
set_unmatched_handler(handler)
Set the given function as the handler that executes when a string is not matched to a configured rule.
The handler receives the UnmatchedText and returns None. It replaces the
default behaviour of raising UnmatchedTextError.
Raises:
| Type | Description |
|---|---|
ScannerConfigurationError
|
if the scanner was created with
|
Source code in src/lectes/scanner/scanner.py
Location
dataclass
Represents where a piece of the scanned text starts.
The offset is 0-based; line and column are 1-based. Columns count
characters, so a tab counts as one column, and only \n starts a new line.
Source code in src/lectes/scanner/models.py
Token
dataclass
Represents a token returned by the scanner.
The scanned token is related to a configuration rule, a string literal and the location in the text where the literal starts.
Source code in src/lectes/scanner/models.py
UnmatchedText
dataclass
Represents a contiguous run of scanned text that no rule matched, and the location where it starts.
Source code in src/lectes/scanner/models.py
ScannerConfigurationError
Bases: ScannerError
The scanner has been configured in a conflicting way.
ScannerError
UnmatchedTextError
Bases: ScannerError
Part of the scanned text does not match any configured rule.
The unmatched text and its location are available as unmatched.