What PyneCore does to your Python before it runs it

PyneCore runs Pine Script logic in Python. The code it runs is called Pyne code, and it looks like ordinary Python. Plain Python would still run it wrong. Before your script runs, PyneCore rewrites it.

This post shows what that rewrite does on a real script. Then it shows how I recently made the rewrite more than twice as fast, and how I checked that its output stayed exactly the same.

Python forgets, Pine remembers

Pine runs your script once for every bar on the chart, oldest bar first. Between two runs it remembers things: the previous values of a series, a counter you keep, the internal state of ta.sma().

A Python function forgets. Call it twice and the second call knows nothing about the first. Its local variables are created when the call starts and thrown away when it ends.

Here is a small Pyne code script that counts how many times a fast moving average crossed above a slow one:

"""
@pyne
"""
from pynecore.lib import script, close, ta, plot
from pynecore.types import Series, Persistent


@script.indicator("Demo")
def main():
    fast: Series[float] = ta.sma(close, 10)
    slow: Series[float] = ta.sma(close, 30)
    crosses: Persistent[int] = 0
    if fast > slow and fast[1] <= slow[1]:
        crosses += 1
    plot(crosses, "Crosses")

If you run that as plain Python once per bar, three things go wrong:

  • crosses starts again from 0 on every bar, so it never gets past 1.
  • fast[1] should be the value of fast on the previous bar. In plain Python fast is just a number, and you cannot index a number.
  • The two ta.sma() calls each need their own memory of past prices, and plain Python has nowhere to keep it.

The @pyne line in the docstring tells PyneCore to fix all of this before the code runs.

What actually runs

This is what the main() function above turns into. I took it from PyneCore’s debug output, which shows names where the real code has plain slot numbers. Everything else is exactly what runs:

@lib.script.indicator('Demo')
@__attach_layout__(__pyne_slot_layout__['main'])
def main(__state__):
    fast = __state__[__slot·main·fast__].add(lib.ta.sma(__st·__ if (__st·__ := __state__[__slot·main·lib·ta·sma·0__]) is not None else __resolve_slot·__(__state__, __slot·main·lib·ta·sma·0__, lib.ta.sma), lib.close, 10))
    slow = __state__[__slot·main·slow__].add(lib.ta.sma(__st·__ if (__st·__ := __state__[__slot·main·lib·ta·sma·1__]) is not None else __resolve_slot·__(__state__, __slot·main·lib·ta·sma·1__, lib.ta.sma), lib.close, 30))
    if not not 1e-10 < fast - slow and (not not ((__cmp1__ := __state__[__slot·main·fast__][1]) <= (__cmp2__ := __state__[__slot·main·slow__][1]) or 1e-10 >= __cmp1__ - __cmp2__)):
        __state__[__slot·main·crosses__] += 1
    lib.plot.plot(__state__[__slot·main·crosses__], 'Crosses')

Nobody has to read this every day, and it is hard on the eyes. Each piece of it solves one of the problems above, though.

main() got a hidden first parameter, __state__, and that is where the memory lives. It is a plain Python list that PyneCore keeps alive between bars and passes in on every call. Each thing the script must remember gets its own position in that list, called a slot.

crosses now lives in a slot. Every read and write of it goes to __state__[__slot·main·crosses__], and because the list survives between bars, the counter keeps counting.

The slots of fast and slow hold a series buffer. On each bar .add(...) pushes the new value into it, and fast[1] became a read from that buffer, where the previous bar’s value is.

The two ta.sma() calls got two different slots, sma·0 and sma·1. On the first bar the slot is empty, and __resolve_slot·__ creates the state for that call. On later bars the call finds its state already in place. Because of this, two ta.sma() calls on the same line never mix up their numbers, and neither does one call inside a function that you call from two places.

The comparisons changed too. fast > slow turned into 1e-10 < fast - slow. TradingView treats two floats as equal when they are closer than 1e-10, and PyneCore has to agree with it, so every comparison operator is rewritten into a form that applies the same tolerance.

The rest of the module, not shown here, gets a small table that describes all the slots: their order, their starting values and which ones are series.

How the rewrite works

PyneCore works on the structure of the code, using Python’s own parser, and never touches the text directly.

Python can turn source code into a tree that describes its structure, called the abstract syntax tree, or AST. The ast module in the standard library gives you that tree. The condition fast > slow looks like this:

>>> import ast
>>> print(ast.dump(ast.parse('fast > slow').body[0].value, indent=2))
Compare(
  left=Name(id='fast', ctx=Load()),
  ops=[
    Gt()],
  comparators=[
    Name(id='slow', ctx=Load())])

A program can walk this tree, find the nodes it cares about and replace them with new ones. Python then compiles the changed tree into bytecode, as if you had written the new code yourself.

PyneCore does this in 38 steps (atm), one after the other, and each step has a single job. One handles Persistent variables, one handles series, one gives every stateful call its own slot, one makes the comparisons tolerant, one rewrites division so that dividing by zero gives na like in Pine, and so on. Keeping the steps separate keeps the system understandable: when something is wrong with division, there is one place to look.

Why the speed of the rewrite matters

The rewrite runs once, when the script is imported, and the result is cached in a .pyc file like any other Python bytecode. The next run loads the cached file and skips the rewrite. It never runs once per bar.

While you are working on a script, though, it runs every time. You change a line, you run the script again, and the whole script has to be rewritten. The same goes for the first run of a script you have just converted. Those are the moments when you are sitting there waiting for the result.

Where the time went

When I measured where the rewrite spent its time, most of it went into walking the tree, and much less into changing it. Every step walked the whole tree from top to bottom to find the few nodes it cared about, and a script has a lot of nodes. Some steps walked over the same nodes again and again, so their cost grew much faster than the size of the script.

Five changes made the difference:

  • I went through the slow steps and removed the repeated walks.
  • The walk itself got cheaper. Python’s AST nodes also store things a walk never needs to look into, like the text of a name or the value of a number, and the walker now skips those fields.
  • Copying parts of the tree got cheaper. Some steps copy code, for example when a function has to exist in two versions, and the copier now skips fields that are empty anyway.
  • Garbage collection is paused during the rewrite. The rewrite creates many small objects quickly, and Python’s garbage collector kept stopping to inspect them. Now it waits until the rewrite has finished.
  • A step can now be written so that it shares one walk with its neighbours and still stays its own separate piece of code. Three of the expression steps already run this way.

The last one follows the rule I cared about most. The steps stayed separate in the code, each with its own name and a comment that says why it stands where it does in the order. Only the walking was merged, and only where that was safe.

How I knew nothing broke

Every script goes through this rewrite, so changing it is risky. A small mistake could change how some script computes on some bar, and nobody might notice for months.

Before changing anything, I built a recorder. It ran the old version on 1,677 input files and saved exactly what came out for each one: the generated code, the full tree including the line and column of every node, and a fingerprint of the compiled bytecode. The inputs were the compiled scripts of Pyne in the Wild and their libraries, PyneCore’s own test files and its built-in library modules.

After every change, the new version had to produce exactly the same output, byte for byte. The only differences I accepted were in two bookkeeping records that I changed on purpose. They describe a module’s dependencies and interface for the cache.

The full PyneCore test suite also had to pass, and I ran the whole Pyne in the Wild corpus against TradingView’s own results again. All 1,000 scripts gave the same results as before the change.

The numbers

I measured the two released versions, PyneCore 6.10.8 and 6.10.9, on 119 real scripts from the Pyne in the Wild corpus. Each script was rewritten three times, and I kept the fastest time:

6.10.86.10.9
All 119 scripts together16.8 s7.5 s
A typical script (median)70 ms33 ms
The slowest script893 ms485 ms

That is about 2.2 times faster overall. A typical script is now rewritten in about a thirtieth of a second.

The cache got more careful too

While I was in there, I also fixed how the cached result is kept up to date.

A script’s rewrite depends on more than the script itself. If the script imports one of your libraries, the rewrite depends on whether the functions in that library keep state, because a function that keeps state needs a slot at every place it is called. Before, changing a library in a way that changed this did not always make the scripts using it get rewritten. Now the cache records what each script relied on, and a change to any of that makes the script get rewritten.

PyneCore also used to keep what it knows about a library in a separate .pynetypes.json file next to your files. That information now lives inside the library’s own .pyc file, so nothing extra shows up in your folders.

If you want the details

This post leaves a lot out. The AST transformation page of the PyneCore documentation describes every step, the order they run in and the reason for that order. To see what your own script turns into, run it with the PYNE_AST_DEBUG=1 environment variable after a change, and PyneCore prints the rewritten code.