Skip to content

io.ErrUnexpectedEOF causes permanent connection pool poisoning via stuck inProgress flag #1298

Description

@henhouse

Summary

A single io.ErrUnexpectedEOF during row reading can permanently poison a pooled connection. After the error, every subsequent query on that connection will fail with:
pq: there is already a query being processed on this connection

The connection is never evicted from the pool because database/sql doesn't see driver.ErrBadConn.

Root cause

PR #1272 added an inProgress atomic flag that is set at the start of query()/Exec() and only cleared when ReadyForQuery is received from the server. If a network error prevents ReadyForQuery from arriving, the flag stays stuck at true.

handleError classifies io.EOF as driver.ErrBadConn, but not io.ErrUnexpectedEOF. When recvMessage does an io.ReadFull on a partially-received message body and the connection drops mid-read, the result is io.ErrUnexpectedEOF, not io.EOF. Since handleError doesn't recognize this, cn.err is never set, IsValid() returns true, and database/sql keeps handing out the broken connection. The CompareAndSwap guard rejects every query, and the non-ErrBadConn error means database/sql won't retry on a fresh connection.

How to reproduce

I've hosted a reproduceable demo of this bug at: https://github.com/henhouse/pq-bug-demo.

It uses a TCP fault-injection proxy between a Go client and a real PostgreSQL 16 instance (via Docker). The proxy forwards all traffic transparently, and then on a specific query it will truncate a DataRow response mid-body -- sending the full 5-byte header but only a fraction of the body bytes before closing the connection. This forces io.ReadFull to return io.ErrUnexpectedEOF.

git clone https://github.com/henhouse/pq-bug-demo
cd pq-bug-demo
make up          # start PostgreSQL 16 via Docker
make test-buggy  # demonstrates the poisoning on v1.12.0
make down

The Makefile also includes make test-downgrade (v1.10.9, no bug) and make test-fix (v1.12.0 + proposed patch) for comparison.

Real-world impact

We hit this in production with CockroachDB, where node drains and rebalances routinely produce partial TCP reads. A single truncated response poisoned the connection pool.

PR Fix

#1299

Activity

  1. added a commit that references this issue on Apr 1, 2026
    6cc34d3
  2. added a commit that references this issue on Apr 1, 2026
    6d77ced
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions