sun::parsing::Lexer
class · Source (opens in a new tab)
class sun::parsing::LexerCreates a token scanner reading from the supplied input stream.
Public Functions
- Lexer
- emitComments
- getNextToken
- getPosition
- getSourceLine
- getSourceText
- getStaticFullRegex
- getTokenDFA
- operator=
- resetInput
- setEmitComments
- setPosition
- ~Lexer
Lexer
Lexer(Lexer &&) noexcept=default
public · function · Source (opens in a new tab)
sun::parsing::Lexer::Lexer(Lexer &&) noexcept=defaultCreates a token scanner reading from the supplied input stream.
Lexer(std::istream &in)
public · function · Source (opens in a new tab)
sun::parsing::Lexer::Lexer(std::istream &in)Creates a token scanner reading from the supplied input stream.
Lexer(const Lexer &)=delete
public · function · Source (opens in a new tab)
sun::parsing::Lexer::Lexer(const Lexer &)=deleteCopying a Lexer would duplicate a position into a shared stream; move only.
Related: Lexer
emitComments
public · function · Source (opens in a new tab)
bool sun::parsing::Lexer::emitComments() constReports whether scanning retains comments as tokens.
getNextToken
public · function · Source (opens in a new tab)
Token sun::parsing::Lexer::getNextToken()Scans and returns the next token from the input.
Related: Token
getPosition
public · function · Source (opens in a new tab)
sun::support::Position sun::parsing::Lexer::getPosition() constReturns the position stored by this object.
Related: sun::support::Position
getSourceLine
public · function · Source (opens in a new tab)
std::string sun::parsing::Lexer::getSourceLine(int lineNum) constGet a specific line from the source buffer (1-indexed).
getSourceText
public · function · Source (opens in a new tab)
std::string sun::parsing::Lexer::getSourceText(int startOffset, int endOffset) constExtract source text substring from buffer (for storing generic method source).
getStaticFullRegex
public · function · static · Source (opens in a new tab)
static const std::string & sun::parsing::Lexer::getStaticFullRegex()Build the full regex string once (expensive string operations).
Public so tests can determinize the same pattern the lexer uses.
getTokenDFA
public · function · static · Source (opens in a new tab)
static DFA & sun::parsing::Lexer::getTokenDFA()One process-wide token DFA, shared by every Lexer.
Per-scan state is just an int, so nothing here is per-instance; scanning only grows the lazily built transition cache. That mutation is safe because the compiler and the LSP are single-threaded. If that ever changes, make this thread_local or guard DFA::step()'s miss path with a mutex.
Related: DFA, Lexer, DFA::step()
operator=
operator=(Lexer &&) noexcept=default
public · function · Source (opens in a new tab)
Lexer & sun::parsing::Lexer::operator=(Lexer &&) noexcept=defaultTransfers the stored state from another instance during move assignment.
Related: Lexer
operator=(const Lexer &)=delete
public · function · Source (opens in a new tab)
Lexer & sun::parsing::Lexer::operator=(const Lexer &)=deleteDisallows assignment so ownership and object identity cannot be duplicated.
Related: Lexer
resetInput
public · function · Source (opens in a new tab)
void sun::parsing::Lexer::resetInput(std::istream &in)Point the lexer at a new input.
There is no per-lexer scan state beyond the buffer and position; the token DFA is shared and stateless.
Related: DFA
setEmitComments
public · function · Source (opens in a new tab)
void sun::parsing::Lexer::setEmitComments(bool emit)Controls whether the lexer emits comments as tokens.
setPosition
public · function · Source (opens in a new tab)
void sun::parsing::Lexer::setPosition(const sun::support::Position &pos)Updates the position stored by this object.
Related: sun::support::Position
~Lexer
public · function · Source (opens in a new tab)
sun::parsing::Lexer::~Lexer()=defaultDestroys this object and releases its owned members.
Public Fields
kEof
public · variable · static · Source (opens in a new tab)
int sun::parsing::Lexer::kEof = -1No documentation comment.
Private Functions
- advance
- commitPosition
- decodeIntegerDigits
- decodeLiteralBody
- isFloatSuffix
- isIntegerSuffix
- isTokenWhitespace
- literalError
- peekByte
- processStringEscapes
- slurp
- toHex
advance
private · function · Source (opens in a new tab)
int sun::parsing::Lexer::advance()Consume one byte and advance line/column.
Returns kEof at end of input, otherwise the byte value in 0..255. This is the single owner of the lexer's position: nothing else moves currentPos forward.
commitPosition
private · function · Source (opens in a new tab)
void sun::parsing::Lexer::commitPosition(int line, int col, int off)Commit a scan position without a whole-Position copy-assign.
Position carries an optional<std::string> filePath and three optional<int>s that the lexer never sets, and assigning them cost one optional<string> copy-assignment per token. Only the three coordinates actually change.
Related: Position
decodeIntegerDigits
private · function · Source (opens in a new tab)
uint64_t sun::parsing::Lexer::decodeIntegerDigits(const std::string &digits, const sun::support::Position &at, int base=10, size_t start=0) constDecode an integer body while checking separators and overflow.
Related: sun::support::Position
decodeLiteralBody
private · function · Source (opens in a new tab)
uint64_t sun::parsing::Lexer::decodeLiteralBody(std::string_view body, bool isByte, const sun::support::Position &at) constDecode the body of a character literal ('a') or a byte literal (b'a') the text between the quotes, which the token regex has already delimited.
A character literal holds one Unicode scalar value: the source is UTF-8, \xNN reaches U+0000..U+007F, and \u{...} names anything above that. A byte literal holds one byte: the source must be ASCII and \xNN covers 00..FF.
Related: sun::support::Position
isFloatSuffix
private · function · static · Source (opens in a new tab)
static bool sun::parsing::Lexer::isFloatSuffix(std::string_view s)Reports whether a literal suffix denotes a floating-point type.
isIntegerSuffix
private · function · static · Source (opens in a new tab)
static bool sun::parsing::Lexer::isIntegerSuffix(std::string_view s)The valid type suffixes of a numeric literal: one per integer type (21u8), and f32/f64 for floats (1.5f32).
isTokenWhitespace
private · function · static · Source (opens in a new tab)
static bool sun::parsing::Lexer::isTokenWhitespace(int c)Reports whether a byte separates tokens as whitespace.
literalError
private · function · Source (opens in a new tab)
void sun::parsing::Lexer::literalError(const sun::support::Position &at, const std::string &message) constReports a malformed literal at its source position.
Related: sun::support::Position
peekByte
private · function · Source (opens in a new tab)
int sun::parsing::Lexer::peekByte() constReads the next input byte without consuming it.
processStringEscapes
private · function · Source (opens in a new tab)
std::string sun::parsing::Lexer::processStringEscapes(std::string_view raw, const sun::support::Position &at) constProcess escape sequences in regular string literals.
Mirrors InterpolatedStringParser::processEscapes (template strings), with " instead of the template-specific ` and $. The shared core for control-character and backslash escapes comes from sun::parsing::simple.
Related: sun::support::Position, InterpolatedStringParser::processEscapes, sun::parsing::simple
slurp
private · function · Source (opens in a new tab)
void sun::parsing::Lexer::slurp(std::istream &in)Read the entire stream into buffer.
Called once per input. Bulk reads, not istreambuf_iterator: the iterator form goes through the streambuf one character at a time and cost ~20% of total lexing instructions, which is the very per-byte overhead the slurp exists to avoid. istream::read() hands off to sgetn() and memcpys whole chunks.
toHex
private · function · static · Source (opens in a new tab)
static std::string sun::parsing::Lexer::toHex(uint32_t value)Formats a code point as hexadecimal for diagnostic messages.
Private Fields
buffer
private · variable · Source (opens in a new tab)
std::string sun::parsing::Lexer::bufferNo documentation comment.
currentChar
private · variable · Source (opens in a new tab)
int sun::parsing::Lexer::currentChar = ' 'No documentation comment.
currentPos
private · variable · Source (opens in a new tab)
sun::support::Position sun::parsing::Lexer::currentPos {1, 1, 0}No documentation comment.
Related: sun::support::Position
dfa_
private · variable · Source (opens in a new tab)
DFA* sun::parsing::Lexer::dfa_ = &getTokenDFA()No documentation comment.
Related: DFA, getTokenDFA
emitComments_
private · variable · Source (opens in a new tab)
bool sun::parsing::Lexer::emitComments_ = falseNo documentation comment.