Index | Thread | Search

From:
Kirill A. Korinsky <kirill@korins.ky>
Subject:
Re: sys/cnmac: fix command buffer leak on transmit failure
To:
OpenBSD tech <tech@openbsd.org>, Klemens Nanni <kn@openbsd.org>, visa@hankala.org
Date:
Mon, 13 Jul 2026 02:40:54 +0200

Download raw body.

Thread
On Fri, 10 Jul 2026 13:18:13 +0200,
Kirill A. Korinsky <kirill@korins.ky> wrote:
> 
> On Thu, 09 Jul 2026 17:55:03 +0200,
> Kirill A. Korinsky <kirill@korins.ky> wrote:
> > 
> > On Thu, 09 Jul 2026 16:22:51 +0200,
> > Visa Hankala <visa@hankala.org> wrote:
> > > 
> > > On Thu, Jul 09, 2026 at 12:28:47AM +0200, Kirill A. Korinsky wrote:
> > > > visa@,
> > > > 
> > > > my ER-4 whcih I used as router had reached interesting state: it floods
> > > > errors like this:
> > > > 
> > > > Jun 30 11:11:39 gw /bsd: cnmac2: cannot allocate command buffer from free pool allocator
> > > > Jun 30 11:11:39 gw /bsd: cnmac2: failed to transmit packet
> > > > Jun 30 11:11:40 gw /bsd: cnmac2: cannot allocate command buffer from free pool allocator
> > > > Jun 30 11:11:40 gw /bsd: cnmac2: failed to transmit packet
> > > > 
> > > > tons of them, I have no idea what had happened before I discovered it in
> > > > that state, but device was ok via serial, but network was dead.
> > > > 
> > > > After reading code near that error I think I had spotted a leak.
> > > 
> > > It seems to me that there is no leak in the code. The local cmdptr
> > > pointer is derived using cmdptr_idx and discarded when the function
> > > returns. cnmac_send_cmd() updates cmdptr_idx only if it was able to
> > > add the command words to the queue.
> > > 
> > 
> > Do you have any idea why it had exploded?
> > 
> 
> What I had found in logs so far.
> 
> That device uses 3 cnamc:
>  - cnmac1 is heavy used to route vlans traffic
>  - cnmac2 is wan port connected directly to modem
>  - cnmac3 is backup wan connected to the same switch to which cnmac1 is
>    connected
> 
> Usually no traffic via cnmac3 until ifstated tries to switch to use it as
> main wan interface.
> 
> I don't have whole logs, just between 7 am and about 11 am where I rebooted
> it, but that log states that "cannot allocate command buffer from free pool
> allocator" had following distribution:
>  - cnmac1 -- 0 records;
>  - cnmac2 -- 3108 records;
>  - cnmac3 -- 14590 records.
> 
> Next, I do have logs from switch but only for cnmac3 connected port (have no
> idea why), and it states:
>  - Jun 30 at 12:25 AM device was disconnected;
>  - Jun 30 at 11:17 AM device was connected again.
> 
> Unfortently I haven't tried to disconnect port physically and I do not
> rememeber does a led lighted up.
> 
> So far it smells that something happened and it consumed all PKO queue.
> 
> But I do not understand, yet, how can it be and how can it lead to that
> state which I saw.
> 

I have added kn@ to the loop because I recall what he had reported in bugs@
the same issue with his device.

I think fix is increase OCTEON_POOL_NELEMS_CMD

OpenBSD currently provides 32 command buffers, each containing 127 usable 64
bit words. A packet command occupies two words, while the last word links
the next buffer; one buffer therefore accommodates 63 packet commands. Every
configured PKO queue owns a current tail buffer.

Each interface permits 1023 unfinished transmissions or uses 17 command
buffers.

So, two "stopped" interface queue requires 34 buffers, what is larger than
32 as a global limit.

My device with 4 interfaces seems need 68 buffers as safe limit. kn@'s with
6 102.

Probably we need to bump that OCTEON_POOL_NELEMS_CMD to 128 or 256. The last
one is safe for 16 interfaces.

visa@, am I wrong?


-- 
wbr, Kirill