Summary
A single io.ErrUnexpectedEOF during row reading can permanently poison a pooled connection. After the error, every subsequent query on that connection will fail with:
pq: there is already a query being processed on this connection
The connection is never evicted from the pool because database/sql doesn't see driver.ErrBadConn.
Root cause
PR #1272 added an inProgress atomic flag that is set at the start of query()/Exec() and only cleared when ReadyForQuery is received from the server. If a network error prevents ReadyForQuery from arriving, the flag stays stuck at true.
handleError classifies io.EOF as driver.ErrBadConn, but not io.ErrUnexpectedEOF. When recvMessage does an io.ReadFull on a partially-received message body and the connection drops mid-read, the result is io.ErrUnexpectedEOF, not io.EOF. Since handleError doesn't recognize this, cn.err is never set, IsValid() returns true, and database/sql keeps handing out the broken connection. The CompareAndSwap guard rejects every query, and the non-ErrBadConn error means database/sql won't retry on a fresh connection.
How to reproduce
I've hosted a reproduceable demo of this bug at: https://github.com/henhouse/pq-bug-demo.
It uses a TCP fault-injection proxy between a Go client and a real PostgreSQL 16 instance (via Docker). The proxy forwards all traffic transparently, and then on a specific query it will truncate a DataRow response mid-body -- sending the full 5-byte header but only a fraction of the body bytes before closing the connection. This forces io.ReadFull to return io.ErrUnexpectedEOF.
git clone https://github.com/henhouse/pq-bug-demo
cd pq-bug-demo
make up # start PostgreSQL 16 via Docker
make test-buggy # demonstrates the poisoning on v1.12.0
make down
The Makefile also includes make test-downgrade (v1.10.9, no bug) and make test-fix (v1.12.0 + proposed patch) for comparison.
Real-world impact
We hit this in production with CockroachDB, where node drains and rebalances routinely produce partial TCP reads. A single truncated response poisoned the connection pool.
PR Fix
#1299
Summary
A single
io.ErrUnexpectedEOFduring row reading can permanently poison a pooled connection. After the error, every subsequent query on that connection will fail with:pq: there is already a query being processed on this connectionThe connection is never evicted from the pool because
database/sqldoesn't seedriver.ErrBadConn.Root cause
PR #1272 added an
inProgressatomic flag that is set at the start ofquery()/Exec()and only cleared whenReadyForQueryis received from the server. If a network error preventsReadyForQueryfrom arriving, the flag stays stuck attrue.handleErrorclassifiesio.EOFasdriver.ErrBadConn, but notio.ErrUnexpectedEOF. WhenrecvMessagedoes anio.ReadFullon a partially-received message body and the connection drops mid-read, the result isio.ErrUnexpectedEOF, notio.EOF. SincehandleErrordoesn't recognize this,cn.erris never set,IsValid()returnstrue, anddatabase/sqlkeeps handing out the broken connection. The CompareAndSwap guard rejects every query, and the non-ErrBadConnerror meansdatabase/sqlwon't retry on a fresh connection.How to reproduce
I've hosted a reproduceable demo of this bug at: https://github.com/henhouse/pq-bug-demo.
It uses a TCP fault-injection proxy between a Go client and a real PostgreSQL 16 instance (via Docker). The proxy forwards all traffic transparently, and then on a specific query it will truncate a DataRow response mid-body -- sending the full 5-byte header but only a fraction of the body bytes before closing the connection. This forces
io.ReadFullto returnio.ErrUnexpectedEOF.The Makefile also includes make test-downgrade (v1.10.9, no bug) and make test-fix (v1.12.0 + proposed patch) for comparison.
Real-world impact
We hit this in production with CockroachDB, where node drains and rebalances routinely produce partial TCP reads. A single truncated response poisoned the connection pool.
PR Fix
#1299