Major slowdown after firmware update

I have an H7 Plus and just updated the firmware on it for the 1st time in about a year. Afterwards my application is running at about 1/4 the speed that it did prior to the update. The code below was processing in the 30-50ms range, now it is taking 100-180ms, which is way too slow for this application.

sensor.reset()  # Reset and initialize the sensor.
sensor.set_pixformat(sensor.GRAYSCALE)  # Set pixel format to RGB565 (or GRAYSCALE)
sensor.set_framesize(sensor.SVGA)  # Set frame size to SVGA (800x600)
sensor.skip_frames(time=500)  # Wait for settings take effect.



    start_time = time.ticks_ms()
    img = sensor.snapshot()  # Take a picture and return the image.
    c = []  # list of average pixel shades in rows 100 through 500
    d = 0  # variable that is offset from a by 101 to align with the entries in list c
    last_sheet = 0  # holds row value of last sheet divide found
    eos = 0  # end of sheet flag
    sheets = []  # holds the row values of sheet divides (dark area between sheets)
    gaps = []  # holds the row difference (# of pixels) between divides (sheet heights)
    slope_hold = []
    d_hold = []
    for a in range(90, 590):  # vertically process image pixels from row 90 through 590
        if len(sheets) < (rec_code+2):
            b = img.get_pixel([330, a])
            b = (b + img.get_pixel([380, a]))
            b = (b + img.get_pixel([420, a]))/3  # use the average pixel shade between 3 columns (380&400&420)
            c.append(b)  # build list c of average pixel shades
            d = a - 101  # offset by 101 to align with entries in list c
            if d >= 10:  # don't start calculations until after 10 to allow for looking back at averages
                slope_a = (c[d-2]-c[d-3])
                slope_b = (c[d-1]-c[d-2])
                slope_c = (c[d]-c[d-1])
                slope = (slope_a+slope_b+slope_c)/3
                slope_hold.append(slope)
                d_hold.append(d)
                if slope < -2 and eos == 0:  # signifcantly darker trend
                    if c[d] > 125:  # make sure in a high white range, else unlikely to be a negative slope of an sheet
                        new_sheet = (a+2)
                        sheets.append(new_sheet)  # insert row # since identified as being significantly darker
                        eos = 1  # set end-of-sheet flag since we think we found a sheet divide
                        if last_sheet == 0:
                            last_sheet = (new_sheet)  # 1st sheet divide found
                        else:
                            gaps.append(new_sheet-last_sheet)  # insert sheet height
                            last_sheet = (new_sheet)
                else:
                    if eos == 1:
                        if slope > 2:  # significantly lighter trend to reset eos flag
                            eos = 0
        else:
            break

    print(time.ticks_diff(time.ticks_ms(), start_time))

Cluade has better code for you:

from ulab import numpy as np

Y0, Y1 = 90, 590
N = Y1 - Y0
COLS = (330, 380, 420)

img = sensor.snapshot()
start_time = time.ticks_ms()

w = img.width()
buf = np.frombuffer(img.bytearray(), dtype=np.uint8)   # no copy

c = np.zeros(N)
for x in COLS:
    c += np.array(buf[Y0 * w + x:Y1 * w + x:w], dtype=np.float)
c /= len(COLS)

slope = (c[3:] - c[:-3]) / 3          # slope[i] is the slope at d = i + 3

cl, sl = c.tolist(), slope.tolist()   # list indexing beats ndarray indexing
limit = rec_code + 2
sheets, gaps, last_sheet, eos = [], [], 0, 0

for d in range(10, N - 11):
    s = sl[d - 3]
    if eos == 0:
        if s < -2 and cl[d] > 125:
            new_sheet = d + 103       # same offset the original arrived at
            sheets.append(new_sheet)
            eos = 1
            if last_sheet == 0:
                last_sheet = new_sheet
            else:
                gaps.append(new_sheet - last_sheet)
                last_sheet = new_sheet
            if len(sheets) >= limit:
                break
    elif s > 2:
        eos = 0

print(time.ticks_diff(time.ticks_ms(), start_time))

From Claude:

I checked equivalence rather than eyeballing it: 1500 randomized synthetic frames through both the original loop and this version, sheets and gaps identical every time, across several rec_code values.

  • indexing has a latent bug that I deliberately preserved. d = a - 101 but c is indexed from a - 90, so the slope is evaluated 11 rows behind the pixel just read, and then the hit is recorded at a + 2 — 13 rows past the sample it fired on. If that offset was meant to be a lag compensation it works by accident; if not, d = a - 90 and new_sheet = d is what he wants.
  • Averaging more columns is now free — widen COLS, or slice a contiguous block and np.mean(..., axis=1). The cost is in the strided reads, not the arithmetic.

Claude is correct that I had an error in changing the starting row in one place but not in another, but apparently he found that it miraculously did not change the outcome. As I understand it, he is pointing out the option to use np.mean as as “free” way to look at more points then just the 3 that I am averaging. BUT unless I’m missing something, neither of these changes would significantly speed up the processing time, nor explain why that time increase 4-fold after the firmware update.

You are currently constructing a large data structure using mallocs. On the H7 Plus, about 2 years ago or more, we moved the heap from internal-only to SDRAM. This means your code, which does a ton of heap mallocs, now has to search SDRAM for free space. While there is more RAM available for larger data structures, it also takes longer to alloc.

.append() of a bunch of small numbers, basically. You could also pre-allocate the list size, and that would speed things up too.

Understood, but why did the firmware upgrade cause the significant slowdown? It was running at a very acceptable speed prior. Is there a way to revert back to the prior version?

Hi Mike,

As mentioned, we extended the heap into SDRAM. Since your code allocates Python objects heavily in a loop, this will slow it down because the time cost per allocation has increased.

You can use your current code mostly unmodified if you simply preallocate the lists to the maximum length they will ever be. Then just assign values to a index in the list and keep a simple valid length counter. This will reduce the malloc churn considerably.

If you want to revert the firmware version you can, however, do you know what version it was? We have old firmware version releases availble on our github.

I purchased the H7 Plus on 5/15/25 and as I recall it updated the firmware the first time I plugged it in, as while I was working on this project last summer I was never ask to update the firmware. If the SDRAM change was made over 2 years ago though, that shouldn’t be the culprit of the slow down. Either way, I’ll try the preallocation, and if that doesn’t work try older firmware and see if one of those options gets me back to the reduced processing time.

You were only using get_pixel from the image module. That function hasn’t really changed. So, the only thing that explains it is the heap moving around. Note that the firmware on the H7 Plus units could be much older than 5/15/25 if it hasn’t been updated.