Thanks to visit codestin.com
Credit goes to github.com

Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
17 changes: 17 additions & 0 deletions InternalDocs/parser.md
Original file line number Diff line number Diff line change
Expand Up @@ -563,6 +563,23 @@ in the generated C parse code that allows to measure how much each rule uses
memoization (check the [`Parser/pegen.c`](../Parser/pegen.c)
file for more information) but it needs to be manually activated.

The C generator also reuses memoized prefixes within consecutive alternatives.
For example, in `prefix ':' NAME | prefix ':' NUMBER`, failure after the first
`':'` normally requires another call to `prefix` and another memo lookup. The
generated code can keep the result and ending position in local variables and
reuse them when trying the next alternative.

This applies only when the shared first item is a memoized rule, including a
left-recursion leader, that the generator can prove consumes input on success.
The locals are reset on each rule-body invocation, including each seed-growing
iteration. Alternative order, cuts, and suffix backtracking are preserved. When
`call_invalid_rules` is enabled, the generated code uses the original rule calls.

The consumption analysis follows grammar items; it cannot inspect arbitrary C
actions. As with memoization, actions must not invalidate cached results. In
particular, suffix actions must not move the parser before their starting mark,
rewrite buffered input, or replace memo entries for earlier positions.

Automatic variables
-------------------

Expand Down
29 changes: 29 additions & 0 deletions Lib/test/test_peg_generator/test_c_parser.py
Original file line number Diff line number Diff line change
Expand Up @@ -145,6 +145,35 @@ def run_test(self, grammar_source, test_source):
TEST_TEMPLATE.format(extension_path=self.tmp_path, test_source=test_source),
)

def test_prefix_reuses_position(self) -> None:
grammar_source = """
start:
| prefix ':' NAME NEWLINE? ENDMARKER
| prefix ':' NUMBER NEWLINE? ENDMARKER
| prefix '=' NUMBER NEWLINE? ENDMARKER
prefix (memo): NAME NAME
"""
self.run_test(grammar_source, """
self.check_input_strings_for_grammar(
valid_cases=['one two : name', 'one two : 3', 'one two = 3'],
invalid_cases=['one = 3', 'one two = name', 'one two :'],
)
""")

def test_prefix_respects_cut(self) -> None:
grammar_source = """
start:
| prefix ':' ~ NAME NEWLINE? ENDMARKER
| prefix ':' NUMBER NEWLINE? ENDMARKER
prefix (memo): NAME NAME
"""
self.run_test(grammar_source, """
self.check_input_strings_for_grammar(
valid_cases=['one two : name'],
invalid_cases=['one two : 3'],
)
""")

def test_c_parser(self) -> None:
grammar_source = """
start[mod_ty]: a[asdl_stmt_seq*]=stmt* $ { _PyAST_Module(a, NULL, p->arena) }
Expand Down
32 changes: 32 additions & 0 deletions Lib/test/test_peg_generator/test_prefix.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
import unittest

from test import test_tools

with test_tools.imports_under_tool("peg_generator"):
from pegen.c_generator import consuming_rules
from pegen.testutil import GrammarParser, parse_string


class ConsumingRuleTests(unittest.TestCase):
def test_predicates_cuts_and_nullable_repeats(self):
grammar = parse_string("""
start: NAME ENDMARKER
positive: &NAME
negative: !NAME
cut: ~ { _PyPegen_dummy_name(p) }
optional: [NAME]
empty_repeat: NAME*
nullable_repeat: optional+
consuming_repeat: NAME+
""", GrammarParser)
self.assertEqual(consuming_rules(grammar.rules), {'start', 'consuming_repeat'})

def test_fixed_point_and_mixed_alternatives(self):
grammar = parse_string("""
start: expression ENDMARKER
expression: expression '+' term | term
term: atom
atom: NAME | '(' expression ')'
nullable: NAME | &NAME
""", GrammarParser)
self.assertEqual(consuming_rules(grammar.rules), {'start', 'expression', 'term', 'atom'})
Original file line number Diff line number Diff line change
@@ -0,0 +1,2 @@
Speed up parsing by reusing memoized rule prefixes across consecutive grammar
alternatives, avoiding repeated rule calls and cache lookups.
Loading
Loading