Repository navigation
[fix](remote-doris) Run a remote Doris scan's query when the coordinator dispatches the plan, not while planning - #68353
morningman wants to merge 10 commits into
Conversation
|
Thank you for your contribution to Apache Doris. Please clearly describe your PR:
|
|
run buildall |
789cd83 to
aefca17
Compare
|
run buildall |
TPC-H: Total hot run time: 27542 ms |
TPC-DS: Total hot run time: 152352 ms |
ClickBench: Total hot run time: 24.15 s |
|
run buildall |
TPC-H: Total hot run time: 27350 ms |
TPC-DS: Total hot run time: 152869 ms |
ClickBench: Total hot run time: 24.19 s |
|
/review |
There was a problem hiding this comment.
Request changes: two P1 correctness gaps remain.
Findings
- Streaming-task cancellation does not give the new statement cleanup a post-quiescence owner. STOPPED skips it, and PAUSED can run it before planning finishes, so the pre-rewrite remote Doris Flight SQL session can still leak until remote
wait_timeout. - The process-wide Awaitility setting removes uncaught raw-worker failures from every suite verdict before the framework has owner-local propagation, allowing product failures to turn green.
Critical checkpoint conclusions
- Goal and proof: The PR closes remote Doris sessions on several previously missed statement paths, but does not accomplish that goal for canceled streaming attempts. The direct cleanup test does not exercise STOPPED or cancellation racing planning.
- Scope and clarity: The FE cleanup changes are focused; the Awaitility switch is process-wide and therefore much broader than the connection-ownership fix.
- Concurrency: The streaming scheduler task thread races the PAUSE/STOP control thread over
ctxandstmtExecutor; the job lock does not quiescebefore(). Regression suites also use raw worker threads;join()provides ordering but no throwable propagation. Connection-map add/drain/remove synchronization otherwise survived the interleaving review, with heavy closes outside its monitor. No new lock-order or deadlock defect was found. - Lifecycle: Prepared execution closes the correct statement generation, Arrow-deferred scans are handed to the coordinator, AutoClose owners finish synchronous use before teardown, and legacy/Nereids coordinators now attempt every scan-node stop. The remaining broken lifecycle is the canceled streaming attempt described inline. No static-initialization issue applies.
- Configuration and compatibility: No configuration, wire/storage format, symbol, rolling-upgrade, persistence, transaction, data-write, or FE-BE variable-passing change is introduced.
- Parallel and conditional paths: Direct/forwarded prepared execution, normal/Arrow ownership, AutoClose callers, legacy/Nereids coordinators, success/failure/retry, PAUSED/STOPPED, and regression suite/STRAY connection ownership were checked. Per-node cleanup isolation and connection-transfer conditions are sound; the canceled-task cleanup suppression is not sound with late planning.
- Tests and results: Added tests cover positive direct cleanup, coordinator failure isolation, prepared execution, and connection tracking, but omit the two cancellation paths and owner-local raw-worker failure propagation. This was a static-only review: the review prompt prohibited builds/tests, so author/CI claims were not treated as independent validation and no local test result is claimed.
- Observability: Existing logs are adequate to diagnose cleanup failures, but logging cannot replace deterministic session close or a failing suite verdict.
- Performance: No substantiated CPU, memory, or complexity regression was found in the per-node stop or connection tracking changes.
- User focus: No additional focus was supplied; the entire 13-file change set and relevant upstream/downstream call chains were reviewed.
All other candidates were either duplicates or dismissed with concrete code evidence. Overall state: changes requested pending fixes for both P1 findings and negative coverage for their failure paths.
| // released here rather than left to the remote frontend's wait_timeout. | ||
| StatementContext statementContext = ctx.getStatementContext(); | ||
| if (statementContext != null) { | ||
| statementContext.close(); |
There was a problem hiding this comment.
[P1] Close canceled attempts only after planning has quiesced
This cleanup is not guaranteed to run at a terminal task boundary. STOPPED only calls cancelAllTasks(false), which marks the task CANCELED; StreamingInsertJob.updateJobStatus force-cleans only PAUSED, while AbstractStreamingTask.execute() skips its own cleanup for every canceled task, so this line is never reached. There is also a PAUSED race before stmtExecutor is assigned: the control thread can close and clear the statement while baseCommand.initPlan(...) is still running, after which a RemoteDorisScanNode can register into the already-drained StatementContext and the canceled task again skips a final close. In both cases the pre-rewrite plan has no coordinator to close that Flight SQL session, leaving it until remote wait_timeout. Please make the task thread perform one final close after before()/run() can no longer register nodes for every terminal status, and cover STOPPED plus cancel-during-planning.
| // the handler in place is whichever suite entered an await() last, so the same exception may just | ||
| // as well reach nobody. A thread's failure reaches its suite through Suite.thread() and the future | ||
| // it returns; a thread a suite started itself is the suite's to join and check. | ||
| Awaitility.doNotCatchUncaughtExceptionsByDefault() |
There was a problem hiding this comment.
[P1] Preserve a suite verdict for raw worker failures
This avoids cross-suite misattribution by making uncaught raw-worker failures unable to fail any suite. The tree still has many Thread.start/new Thread workers that run SQL or assertions and whose parents only call join(); join() never rethrows, and the framework installs no per-suite uncaught-error collector. Suite.thread() can propagate through its future, but those existing raw-thread callers do not use it. The manual result in this PR also explicitly shows the new behavior: a failing daemon only prints to stderr and fails no suite. Please add owner-local propagation (for example, migrate these workers to futures plus get(), or add a per-suite collector checked after joins) before disabling the only process-level signal, so product regressions cannot turn green.
|
run buildall |
TPC-H: Total hot run time: 27814 ms |
TPC-DS: Total hot run time: 151973 ms |
ClickBench: Total hot run time: 24.1 s |
FE Regression Coverage ReportIncrement line coverage |
FE UT Coverage ReportIncrement line coverage |
81bc1dd to
7b33695
Compare
|
run buildall |
TPC-H: Total hot run time: 28712 ms |
… split sources are released ### What problem does this PR solve? Issue Number: None Related PR: apache#68338 Problem Summary: StmtExecutor.handleQueryWithRetry dispatches the same plan again when an attempt fails with an RPC error. Before the retry, the failed attempt's cancel() has stopped the scan nodes, and a batch-mode scan's stop() unregisters the split sources its scan ranges point the backends at. So the retry of a query with a batch-mode external scan (one whose backends fetch the splits from the frontend while they scan) could not succeed: its backends asked for a split source that was gone and failed with "Split source <id> is released", which the query reported in place of the original error. apache#68338 added ScanNode.cannotBeRedispatched() so the retry rethrows the original error instead, but only the remote Doris scan answered true. It is now a rule of every scan with a split assignment: an assignment serves one dispatch, so once it is stopped the same plan cannot be dispatched again. An assignment that was never stopped (the attempt failed before anything stopped it) does not block the retry. ### Release note A query with a batch-mode external table scan that fails with an RPC error is no longer retried with the same plan, which could only fail with "Split source ... is released"; it reports the original error. ### Check List (For Author) - Test: Unit Test - ScanNodeDispatchTest: a scan whose split assignment was stopped cannot be redispatched; one not stopped yet, or one without a split assignment, can. - Behavior changed: Yes. The same-plan retry of a query with a batch-mode scan now rethrows the original error instead of failing on a released split source. - Does this need documentation: No Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…its are all generated ### What problem does this PR solve? Issue Number: None Related PR: apache#68338 Problem Summary: A backend fetches the splits of a batch-mode scan from the split source on this frontend (SplitSource.getNextBatch, through FrontendServiceImpl.fetchSplitBatch), asking for up to remote_split_source_batch_size (1000) of them. The fetch takes what is queued for the backend, then polls the queue again with a 100 ms timeout, and returns early only once fetch_splits_max_wait_time_ms (1000) has passed. So the fetch that empties the queue waits another 100 ms on it before it asks the assignment whether it is done - also when the generator has finished and nothing can arrive any more. For a hive / iceberg / paimon batch scan that is a 100 ms tail per backend at the end of the scan. For a scan whose generator queues every split at once and finishes before the backends fetch anything - the remote Doris scan, once it runs its remote query when the coordinator dispatches the plan (a following commit) - it is a fixed 100 ms before the first DoGet of every backend holding an endpoint, on every query, small lookups included. The fetch now takes what is queued without waiting once the assignment needs no more splits: its generator finished, or it was stopped or failed. Every generator queues its splits before it calls finishSchedule() (the partition-batch flavour finishes only after every batch task did), so nothing is lost. A backend holding splits still fetches twice: its second fetch learns, at once, that the source is done. Measured with the remote Doris commit that follows: a one-row lookup over a remote Doris catalog (`SELECT id, v FROM <catalog>.db.t WHERE id = 1`, use_arrow_flight = true, one FE and one BE on a laptop, the catalog reading the same cluster), 30 runs in one session right after an FE restart, took a median 160 ms without this commit and 58 ms with it (46 ms on a warm FE): the 100 ms wait is gone. ### Release note None ### Check List (For Author) - Test: Unit Test, Manual test - SplitAssignmentTest: once the generator has finished, a fetch takes the last splits without waiting on the queue, and the next fetch learns that the source is done. - The latency above, measured on a local cluster. - Behavior changed: No - Does this need documentation: No Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
FE Regression Coverage ReportIncrement line coverage |
7b33695 to
a9535b5
Compare
|
run buildall |
TPC-H: Total hot run time: 29284 ms |
TPC-DS: Total hot run time: 152201 ms |
ClickBench: Total hot run time: 24.49 s |
a9535b5 to
430a606
Compare
|
run buildall |
TPC-H: Total hot run time: 29171 ms |
TPC-DS: Total hot run time: 152300 ms |
ClickBench: Total hot run time: 24.32 s |
FE Regression Coverage ReportIncrement line coverage |
…the connection ### What problem does this PR solve? Issue Number: None Related PR: apache#68338 Problem Summary: A backend keeps up to max_client_cache_size_per_host (10) idle thrift clients per server (ClientCacheHelper) and hands the first one out again for the next call. A frontend restart closes every connection a backend had cached to it, but the cache keeps the clients and hands a dead one out: the call on it fails, with "No more data to read." when the request reached a socket the server had closed, with "write() send(): Broken pipe" afterwards. Thrift does not close a socket a call failed on (TSocket only throws), and release_client puts the client back as it is, so it fails again for the next caller - until a caller that reopens on a transport error happens to draw it. Most backend-to-frontend calls reopen and send again (ThriftRpcHelper, ReportExecStatus). The split fetch of a batch scan (RemoteSplitSourceConnector, fetchSplitBatch) does not, and must not: the frontend hands out the splits of a batch as it answers (SplitSource.getNextBatch polls them off its queue), so a request sent again after an answer was lost would skip that batch and the scan would return fewer rows. The scan fails instead: right after a frontend restart, a batch-mode hive / iceberg / paimon query that frontend runs can fail with "Failed to get batch of split source". The next commit makes every remote Doris query fetch its splits that way, on every candidate backend. A cached client sits idle between calls, so a server that closed its connection shows on the socket before the next call: its end of stream, or an error, is waiting to be read. get_client now peeks at the socket of the cached client it would hand out (ThriftClientImpl::peer_closed, a non-blocking recv(MSG_PEEK | MSG_DONTWAIT)) and closes the client instead, taking the next cached one or connecting anew. Bytes waiting to be read leave the client as it is. reopen_client and get_client share the closing (_close_client). A server host that went away without closing the connection (no FIN, no RST) still looks idle: the call on it fails when its receive times out, as before. ### Release note A query that reads a batch-mode external table no longer fails with "Failed to get batch of split source" right after the frontend running it restarts. ### Check List (For Author) - Test: Unit Test, Manual test - client_cache_test (3 cases, all passing under ASAN): peer_closed tells a connection the server closed from an idle open one; an idle cached client is handed out again, on the same connection; a cached client whose connection the server closed is not handed out, the cache connects anew. With the check in get_client taken out, the last case fails. - Manual, on a local cluster (one FE, one BE) with a remote Doris catalog pointing back at itself: 20 remote Doris queries, then three FE restarts with 20 remote Doris queries right after each - none failed, and the backend closed 9 cached clients whose connection the restarted frontend had closed ("Close a cached client to ... whose connection the server closed" in be.INFO). Earlier local runs of the next commit without this change logged remote Doris queries failing right after a frontend restart with "Failed to get batch of split source: No more data to read." and "...: write() send(): Broken pipe". - Behavior changed: Yes. A backend connects anew instead of using a cached connection its server closed. - Does this need documentation: No Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…tor dispatches the plan, not while planning ### What problem does this PR solve? Issue Number: None Related PR: apache#68338 Problem Summary: A remote Doris scan (a `type = doris` catalog with `use_arrow_flight = true`) opens a Flight SQL session on a remote frontend, runs its query there (GetFlightInfo) and hands the endpoints of the result to the backends, which read them with DoGet. The session is a connection of the catalog user on the remote frontend, and only a CloseSession, a KILL or wait_timeout ends it. apache#68338 made the scan hold it until the coordinator stops the scan. But the scan opened the session, and ran the remote query, while the plan was being translated (getSplits), before any coordinator existed. Many plans never get one, so every one of them needed a patch of its own: - apache#68338 registered the scan with the StatementContext as a fallback owner. That relies on whoever runs the statement closing it: a direct COM_STMT_EXECUTE, a statement under AutoCloseConnectContext (EXPORT, ANALYZE), a streaming insert task and an http_stream load (whose "only a TVF may be scanned" check runs after planning) do not, and each left a session open on the remote frontend until its wait_timeout. An Arrow Flight SQL query had to hand the scan from the statement over to its deferred coordinator. - EXPLAIN was told apart by matching the statement text against "explain". A job's insert task and a streaming insert task plan with no statement text, so the match threw a NullPointerException and neither could read a remote Doris table at all. A comment before EXPLAIN ran the remote query anyway, and when a multi-statement packet could not be split, the SELECT after an EXPLAIN returned no rows. - The remote query ran for plans that are only inspected: the INSERT OVERWRITE probe plan (so INSERT OVERWRITE and every materialized view refresh ran it twice), the plan CREATE JOB validates, the plan a streaming insert task builds to rewrite its TVF, and a statement refused after planning (a SQL block rule on the scan, an INSERT whose label is taken). A query waiting in the workload group's queue held the session and the remote result buffers meanwhile. - A same-plan retry had to be refused with a remote Doris specific rule. The scan now runs its query when the coordinator dispatches the plan, on the split assignment the previous commits introduced: - The scan is always in batch mode and plans without a sample split (needsSampleSplit() false). Planning builds the remote query (convertPredicate; EXPLAIN shows it as before) and one scan range per backend pointing at a split source, and contacts nothing. - startSplit(), which the split assignment calls when the coordinator starts the scan (exec(), once the query is admitted), runs the query on the remote frontends in turn as before, hands the session to the split assignment (closed at once if the scan was stopped meanwhile) and queues the endpoints as the splits. A remote failure fails the dispatch with the same "Failed to execute query" message. - The handshake that opens the session gets a deadline, the catalog's query_timeout_sec, as GetFlightInfo has. It had none, and the coordinator now runs the remote query while it holds the query's admission slot: a remote frontend that accepts the connection but never answers would hold the slot, and the connection thread, for good - neither KILL nor the query's timeout interrupts the wait. A dispatch's remote work is bounded by query_retry_count x (connect + handshake + GetFlightInfo + closing a failed attempt). - The coordinator stops the split assignment when it closes or cancels, which sends the CloseSession, and the split assignment is what keeps a deferred Arrow Flight SQL coordinator alive (ScanNode.coordinatorMustOutliveDispatch), as for any batch scan. - Gone: isExplainStatement(); the session fields of the scan and its stop(), coordinatorMustOutliveDispatch() and cannotBeRedispatched() overrides (the general rule of the previous commit covers it); the StatementContext fallback (stopScanNodeAtClose, handOverScanNodesToDeferredCoordinator, the step in close()) and the hand-over in the deferral gate. What changes besides: every candidate backend runs a scan instance of a remote Doris scan (most of them reading nothing when the remote query returns few endpoints), as for any batch scan, a backend that holds an endpoint fetching its splits from the frontend twice (the second fetch learns at once that there are no more) and any other once; the split count a SQL block rule sees for the scan is 0 rather than the number of endpoints; the time of the remote query moves from the plan time to the schedule time of the profile. ### Release note A remote Doris catalog (use_arrow_flight = true) runs the query on the remote cluster only when the local query runs: EXPLAIN, a statement that fails before it runs and the plans INSERT OVERWRITE and CREATE JOB only inspect no longer reach the remote cluster, and INSERT OVERWRITE runs the remote query once instead of twice. Jobs and streaming insert jobs can read remote Doris tables, which failed with a NullPointerException before. A remote Doris query whose remote frontend accepts the connection but never answers gives up after the catalog's query_timeout_sec instead of waiting for good. ### Check List (For Author) - Test: Unit Test, Regression test - RemoteDorisScanNodeTest: planning contacts no remote frontend and leaves the split sources unreachable; dispatch runs the query once, and the coordinator's cancel and the close that follows end its session once; a failed query closes its sessions and fails the dispatch; a session opened after the scan was stopped is closed at once; a stopped scan tries no remote frontend; a coordinator stops the scan after one that fails to stop; the handshake with a remote frontend that never answers gives up at the timeout. - external_table_p0/remote_doris/test_remote_doris_flight_session: with a SQL block rule refusing every query of the catalog user on the remote frontend, EXPLAIN (also with a comment before it) works, an INSERT with a used label fails on the label and CREATE JOB creates a job reading the remote table; a job's insert task reads the remote table; no Flight SQL session of the catalog user is left after any statement. The rows read are checked against a generated .out. - Behavior changed: Yes. See the release note, and the scan instances on every candidate backend. - Does this need documentation: No Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…d when its plan is never dispatched ### What problem does this PR solve? Issue Number: None Related PR: apache#68338 Problem Summary: A batch scan that plans with its first split (FileQueryScanNode#needsSampleSplit: the hive / iceberg / paimon / MaxCompute batch scans) starts generating its splits while the plan is translated, and the coordinator dispatching the plan stops the generation when it closes or cancels. A plan that is translated and then dropped gets no coordinator, and only EXPLAIN and the INSERT OVERWRITE probe plan stopped what theirs had started. Every other dropped plan left the generation running: - every successful CREATE JOB ... DO INSERT ... SELECT (CreateJobInfo validates the job's statement with initPlan(false) and drops the plan); - every task of a streaming insert job, whose before() plans the job's statement only to rewrite its TVF and drops that plan, and an attempt failing between planning and dispatch; - an INSERT whose label is taken, refused when its transaction begins after planning; - a statement a SQL block rule refuses after planning (checkBlockRulesByScan), and an http_stream load refused after planning because it reads more than its TVF; - the plan an INSERT drops because the target table changed while it was planned, the plan of a DELETE whose predicate reads another table (it falls back to DeleteFromUsingCommand, which plans the statement again), and the attempt a cloud re-plan retries. The streaming generator of an iceberg scan (num_files_in_batch_mode matched files or more, 1024 by default) queues its splits one at a time into a queue of 10000 per backend. Once a backend's queue is full it waits, 100 ms at a time, for as long as the assignment needs more splits: forever, for an assignment nobody stops. So each such statement held one of the 64 threads of the scheduleExecutor that every batch-mode external scan shares, and the splits it had queued, until the frontend restarted; after 64 of them, the split generation of every batch-mode hive / iceberg / paimon scan on that frontend stalls and those queries time out. A streaming insert job reading such a table besides its TVF did that with every task, every max_interval (10 s by default). The statement now owns what its planning started until a coordinator takes it over: - FileQueryScanNode registers the split assignment it starts while planned with the statement - first, for a start that fails half way - and starts it with SplitAssignment#startWhilePlanning. - SplitAssignment#start, which the coordinator calls in exec() (ScanNode.startAll), takes the assignment over: from then on only the coordinator stops it. - Where the statement drops a plan and goes on, it stops that plan's scans (ScanNode.stopAllUndispatched, which logs a failed split generation of the dropped plan as one, rather than throwing it to be logged as a failure to stop; EXPLAIN and the INSERT OVERWRITE probe, which already stopped theirs, use it as well): an INSERT planned again because its target table changed; a DELETE, whose own plan no coordinator ever takes (a delete by predicate reads only its filter, any other falls back to DeleteFromUsingCommand, which plans again); the plan a streaming insert task's before() builds to rewrite the TVF, as soon as it has it. - Where the statement ends, it stops the registered assignments no coordinator took over and forgets all of them (StatementContext#stopUndispatchedSplitAssignments, deciding under the lock the takeover takes): StatementContext#close, for a statement of a connection, a forwarded statement, a job task and an MTMV task; the end of a binary COM_STMT_EXECUTE, which does not close the context its prepared statement keeps until the next execution (MysqlConnectProcessor#handleExecute); the request of an http_stream load. So does the end of an attempt the next one plans again: every attempt of a streaming insert task, on the task thread and a canceled attempt too (AbstractStreamingTask#endAttempt; the streaming scheduler runs the task itself, not through TaskProcessor), and the attempt a cloud re-plan retries (StmtExecutor#queryRetry). An Arrow Flight SQL query whose coordinator outlives its statement had its assignments taken over, so they stay. - Forgetting them matters as much as stopping them: the context of a binary COM_STMT_EXECUTE lives on in its prepared statement, and an assignment holds its scan node, through it the whole plan, and the splits the backends did not fetch. A stopped assignment now also drops those splits (SplitAssignment#stop), so whatever still references it keeps none of them. - A statement can also end without either (a statement under AutoCloseConnectContext failing before dispatch). So the generator of an assignment started while planned stops it itself once a backend's queue is full and no coordinator took it over within the statement's timeout (getExecTimeoutS, counted from planning): the timeout checker kills a statement of a connection older than that, so no coordinator takes the plan any more. Its log names the scan and the query that planned it. - The stop at the end of a statement never throws: a failure of the split generation of a plan no backend read is logged, with the scan and the query, instead of failing the end of the statement. ### Release note Fix that a CREATE JOB, a streaming insert job, or a statement refused or failing after planning, whose plan reads a large iceberg table in batch mode, left a frontend thread generating splits forever; enough of them stalled every batch-mode external table scan on that frontend. ### Check List (For Author) - Test: Unit Test - SplitAssignmentTest: the coordinator takes over an assignment started while planned, and the end of the statement then leaves it alone; one no coordinator took is stopped, without throwing when its generation had failed; once the backend's queue is full, the generator stops an assignment whose plan stayed undispatched past the timeout, but keeps waiting for the backends once a coordinator took it over; a stopped assignment drops the splits it still queued, and queues none afterwards - including the batch a generator waiting on a full queue offers into the room stop() made. - ScanNodeDispatchTest: a dropped plan is stopped without throwing the failure of its split generation. - StatementContextTest: close() stops every registered assignment and forgets it. - MysqlConnectProcessorExecuteEndTest: a binary COM_STMT_EXECUTE leaves the context its prepared statement keeps holding no split assignment, and stops the one of a plan refused before dispatch. - StreamingInsertTaskAttemptEndTest: before() stops the scans of the plan it builds to rewrite the TVF as soon as it has it; the end of an attempt stops what its planning left undispatched, for an attempt canceled by STOP JOB, and for a PAUSE JOB during planning, whose control thread drops the attempt's context before the plan registers. - DeleteFromCommandTest, FrontendServiceImplHttpStreamPlanTest, StmtExecutorReplanRetryTest: the plan of a DELETE is stopped, as a dropped plan, as soon as it is planned; an http_stream load refused after planning, and the attempt a cloud re-plan retries, stop what their planning started. - RemoteDorisScanNodeTest: a scan that plans with its first split is stopped when its statement ends undispatched, and left to the coordinator that dispatched it. - Behavior changed: No - Does this need documentation: No Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…generator assigns it a split ### What problem does this PR solve? Issue Number: None Related PR: apache#68338 Problem Summary: The split assignment of a batch-mode scan keeps a queue of scan range batches per backend: the generator queues into it (appendBatch), the backend's split source takes from it (getAssignedSplits), and whichever comes first creates it. The generator created a queue of 10000 batches, so that a generator getting ahead of a backend waits for it; a backend's fetch created an unbounded one. A backend fetches as soon as its scan starts, so a backend the generator had not assigned a split by then - while the generator was still reading the table's metadata, say - got an unbounded queue, and the generator never waited for that backend: what it got ahead with piled up on the frontend's heap, which is what the bound keeps a scan of millions of files from doing. Pre-existing; found while testing the stop of a split generation whose plan is never dispatched, when a test that fetched before the generator queued its first batch made the generator fill an unbounded queue until the test JVM ran out of heap. Both now create the same bounded queue. ### Release note Fix that the frontend could keep most of the splits of a huge batch-mode external table scan in memory instead of waiting for the backends to fetch them. ### Check List (For Author) - Test: Unit Test - SplitAssignmentTest: the queue a backend's fetch creates before its first split is bounded like the one the generator creates. - Behavior changed: No - Does this need documentation: No Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… Doris accessors once their thread ends, and stop Awaitility blaming the awaiting suite ### What problem does this PR solve? Issue Number: None Related PR: apache#68338 Problem Summary: apache#68338 made SuiteContext record every Doris connection its thread-local accessors open, with the thread that opened it, close those of finished threads on each statement, and close whatever is left when the suite ends. Three gaps remained: - Only getConnection() ran the sweep; a suite whose statements go through the master or the Arrow Flight SQL accessor (the arrow_flight_sql group) never did. - Recording a connection and the final drain were not serialized, so a connection a thread registered while the suite was draining was lost and stayed open. - The final drain closed the connections of threads still running. Some suites leave a thread running on purpose: test_active_queries and test_backend_active_tasks return at once and leave a daemon thread polling their system table for five minutes. Closing its connection under it, or refusing it a new one, failed that thread's next statement. Awaitility, which by default installs every await() as the JVM's default uncaught-exception handler and rethrows from the awaiting thread whatever any thread threw uncaught meanwhile, then handed that failure to an unrelated suite that happened to be awaiting: P0 build 1054369 failed test_partial_update_insert_schema_change that way. - OpenedDorisConnections (new): the table of a suite's connections with the threads that opened them, each kept exactly as long as its thread runs; add and drain are serialized. The drain at the end of the suite closes the connections of finished threads and hands those of running threads to STRAY, the table shared by all suites, where a registration after the drain goes too. Every statement of any suite closes the strays whose thread has finished, and RegressionTest closes the rest after the last run - under the lock of add() as well, a connection registered after that being closed at once. - All three thread-local accessors (getConnection, getMasterConnection, getArrowFlightSqlConnection) run the sweep. - RegressionTest.initGroovyEnv: Awaitility.doNotCatchUncaughtExceptionsByDefault() next to pollInSameThread(). A thread a suite started itself is that suite's to join and check; Suite.thread() reports through its future. ### Release note None ### Check List (For Author) - Test: Unit Test, Manual test - OpenedDorisConnectionsTest. - SuiteContextDorisConnectionsTest: each of the three accessors closes the connections of the threads that have finished; the end of the suite closes those of finished threads and leaves the connection of a thread still running to the strays, which the next statement of another suite closes once the thread has finished. - Locally, two scratch suites (a daemon like test_active_queries's, a neighbour inside Awaitility.await()) fail with the first cut of this change the way the P0 run did and pass with this one, the daemon's connection closed 100ms after its thread ended. - Behavior changed: No - Does this need documentation: No Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… a thread it started ### What problem does this PR solve? Issue Number: None Related PR: apache#68338 Problem Summary: The previous commit stopped Awaitility from installing every await() as the JVM's default uncaught-exception handler, which failed whichever suite happened to be awaiting for any thread's failure. Wrong as that attribution was, it was the only thing that could turn the failure of a raw Thread a suite starts and join()s into a red build: join() never rethrows, 83 suites do this, and the framework ran no code on such a thread. Without it those failures went to stderr and failed nothing. UncaughtThreadFailures is now the JVM's default handler. A thread a suite constructs inherits the suite's collector - an InheritableThreadLocal set on the suite's thread in ScriptContext.createAndRunSuite, through any depth of threads started from it - and its failure is recorded there. A task of Suite.thread() runs as the suite that submitted it: a worker of the run-wide pool keeps the collector of whichever suite's submission constructed it, so the task carries its submitter's, and a thread the task starts inherits that one. Suite.doLazyCheck() throws the first failure once the body and its lazy checks are over, that is after every thread the suite joined has ended; a suite that fails on something else first carries the failures along, suppressed - often the cause, as when a thread whose statement failed leaves the body a result it then fails on. A failure after that (a thread left running past its suite's end, as test_active_queries intends) or on a thread no suite started fails nothing; every one is logged with the thread and suite names. ### Release note None ### Check List (For Author) - Test: Unit Test, Manual test - UncaughtThreadFailuresTest, including that a thread no suite started fails nothing and is logged as such. - SuiteThreadFailuresTest, running suites the way RegressionTest does: a suite whose joined thread dies fails, one whose threads end well passes; a thread a task of Suite.thread() starts fails the suite that submitted the task, not the one whose submission constructed the pooled worker; a thread a lazy check starts after the body returned fails the suite; a suite failing on something else carries what its thread died of, suppressed. - Four scratch suites against a local cluster with -suiteParallel 4: the suite whose raw joined thread fails a statement now fails, with "Thread Thread-3 of suite joined_worker_fails died with an uncaught exception"; the suite whose joined thread (and a thread started from it) runs fine passes; the suite that returns at once and leaves a daemon whose statement fails 2s later passes, the failure logged as "started by suite stray_daemon whose verdict is already taken; it fails no suite"; and the suite spending those seconds inside Awaitility.await() - the one P0 build 1054369 blamed - passes. - Behavior changed: No - Does this need documentation: No Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ing it finish ### What problem does this PR solve? Issue Number: None Related PR: apache#66331 Problem Summary: An External regression run lost its FE in the middle of the run. fe.out ended with `SIGSEGV ... Problematic frame: C [libadbc_driver_sqlite.so+0xc308d] sqlite3FindTable+0xcd`, and every suite after that failed with "Connection refused". The FE died within about 120 ms of `test_adbc_sqlite_catalog_scan` dropping its catalogs. A query on one of those catalogs had just started three column-statistics loads on the STATS_FETCH pool. For an ADBC table such a load ends in `PluginDrivenExternalTable.getColumnStatistic` -> `AdbcConnectorMetadata.getTableHandle` -> `getObjects`, which runs SQL in the SQLite driver. With `meta.cache.adbc.metadata.enable=false` it does so on every call. One of the three loads never logged a result. Root cause: DROP CATALOG (and ALTER CATALOG, which resets the catalog) closes the connector on its own thread, and `AdbcClient.close()` closed the native `AdbcDatabase` at once. In the JNI driver, `JniDatabase.close()` first closes every connection still open on the database. It does that from the closing thread, while other threads may still be inside the driver on those connections. Neither the JNI bridge nor the C driver guards against this, so the native connection (and SQLite's schema) is freed under a thread that is still reading it. Arrow memory had the same problem: the allocator was closed under readers that still held buffers ("Memory was leaked by query"). Reproduced outside FE with the plugin's real `AdbcClient`: four threads loop over `getObjects` + `getTableSchema` through `withConnection` against the real SQLite driver, and the main thread closes the client. Before: 8 of 8 runs crashed the JVM within 50 closes. Two crashed in exactly `sqlite3FindTable` <- `sqlite3LocateTable` <- `selectExpander`; the rest crashed in `AdbcConnectionGetObjects` / `AdbcConnectionGetTableSchema` / `PrivateArrowSchemaDeepCopy`, at a null pc, or with SIGABRT. After: 8 runs x 200 closes (about 113k calls) all survived, with no allocator leak. Fix: `AdbcClient` counts the calls inside `withConnection`. `close()` now only refuses new calls. It releases the database and the allocator itself when no call is in flight. Otherwise the last call to leave releases them. A failure of that deferred release is logged with the driver's path, as the catalog logs a failed connector close, rather than thrown at a caller that did not close anything: that call's own work succeeded. `close()` does not wait for the calls in flight, so DROP/ALTER CATALOG are not held up by a slow remote source. ### Release note Fix an FE crash (SIGSEGV in the ADBC driver) when an ADBC catalog is dropped or altered while a statement or a statistics load is still reading it. ### Check List (For Author) - Test: Unit Test - AdbcClientTest +4 (run with the real JNI bridge and SQLite driver, 0 skipped): a call that is still on its connection when the client is closed can keep using it (real driver); the last call to leave a closed client releases the database exactly once and only after it leaves; an idle client releases at once; a release that fails as the last call leaves is attempted once and does not fail that call, which returns its own result. With the old close logic put back, the first two fail ("Native connection handle is closed", "the database was released under an open connection"); with the deferred release's failure thrown instead of logged, the fourth does. - Manual test: the stress program above, before and after. - Checkstyle: 0 violations in fe-connector-adbc. - Behavior changed: No (DROP/ALTER CATALOG still return without waiting; the native driver of an ADBC catalog is now freed when its last in-flight call finishes instead of under it) - Does this need documentation: No Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
430a606 to
bdedd76
Compare
|
run buildall |
TPC-H: Total hot run time: 28934 ms |
TPC-DS: Total hot run time: 151851 ms |
ClickBench: Total hot run time: 24.46 s |
BE UT Coverage ReportIncrement line coverage Increment coverage report
|
FE UT Coverage ReportIncrement line coverage |
BE Regression && UT Coverage ReportIncrement line coverage Increment coverage report
|
FE Regression Coverage ReportIncrement line coverage |
What problem does this PR solve?
Issue Number: None
Related PR: #68338 (this is its follow-up), #68101 / #68266 (why a leaked Flight SQL session costs the catalog user's connection quota on the remote FE), #68743 (the same planning-time execution problem in the ADBC catalog)
Problem Summary:
In short. A remote Doris scan ran its query on the remote cluster, and opened the Flight SQL session it runs in, while the local plan was being translated - before any coordinator existed. Every path that plans without running then needed a patch of its own to end that session, and some had none. A few could not read remote Doris at all (jobs, streaming jobs), and some ran the remote query for nothing (EXPLAIN with a leading comment, the INSERT OVERWRITE probe plan, CREATE JOB). This PR makes the scan run its query only when the coordinator dispatches the plan. The query runs when the coordinator starts the scan's split assignment in
exec(), and that split assignment owns the session until the coordinator stops it on close or cancel. A plan that never runs never reaches the remote cluster, and the session has exactly one owner. The first version of this PR patched three more statement owners; that approach is replaced by this one.The remote Doris change builds on the split assignment of the batch-mode hive / iceberg / paimon / MaxCompute scans, and fixes five things there on the way: a same-plan retry that could only fail (item 2), a 100 ms wait at the end of every backend's fetches (item 3), a backend's split fetch on a connection the frontend had closed, which failed the query right after a frontend restart (item 4), a dropped plan's split generation that held a shared frontend thread forever (item 6), and an unbounded split queue (item 7). Two regression-framework fixes (items 8 and 9) and an unrelated ADBC crash fix (item 10, which can be picked on its own) ride along.
Background. A
type = doriscatalog withuse_arrow_flight = truereads another Doris cluster over Arrow Flight SQL.RemoteDorisScanNodeopens a Flight SQL session on a remote FE (the handshake), runs a query there (GetFlightInfo), and hands the endpoints of its result to the local BEs, which read them from the remote BEs with DoGet. The session is a connection of the catalog user on the remote FE. It counts againstmax_user_connectionsthere, and only CloseSession, KILL orwait_timeout(8h) ends it. It cannot be closed as soon as GetFlightInfo returns, because the remote FE cancels what a closed session still runs. #68338 therefore made the scan hold the session until the coordinator stops it. Because the session was opened while planning, it also registered the scan with theStatementContextas a fallback owner for plans that never get a coordinator.The problem, and what it cost
EXPLAIN SELECT ..."explain"/* comment */ EXPLAIN SELECT ...EXPLAIN ...; SELECT ...NullPointerExceptionin the EXPLAIN match, cannot read remote Doris at allWITH LABELalready used, txn limit)before()(TVF rewrite plan)COM_STMT_EXECUTE,AutoCloseConnectContext(EXPORT, ANALYZE), streaming task, http_stream refused after planningwait_timeouthandleQueryWithRetry)Split source N is releasedquery_timeout_sec, and the next remote frontend is triedFailed to get batch of split sourceHow this PR fixes it
[refactor](fe)A split assignment owns what a batch scan holds for the backends, from dispatch to coordinator close.SplitAssignmentgets a start/stop pair.start()registers the split sources withSplitSourceManagerand starts the generator.stop()stops the generator, unregisters the sources, and closes what was handed over throughaddCloseable(); it rethrows a generation failure only after all of that is released.stop()is closed at once, and a stopped assignment does not start.ScanNode.start()/stop()drive the assignment. Both coordinators callScanNode.startAllinexec(), once the query is admitted and before the fragments are sent, andScanNode.stopAllonclose()/cancel().stopAllkeeps going past a node that fails to stop.FileQueryScanNode.needsSampleSplit()defaults to true. A scan that plans with its first split (its location type sets the scan params) still starts generating while planning - item 6 makes the statement own that generation until a coordinator takes it over. A scan that answers false generates nothing until dispatch; itsstartSplithas to queue its splits before it returns, sincestart()waits only until the first split is published.[fix](fe)ScanNode.cannotBeRedispatched()becomes a general rule: a stopped split assignment cannot serve another dispatch. The same-plan retry of a batch-mode external scan now reports the original error instead ofSplit source N is released.[improvement](fe)A backend's split fetch no longer waits 100 ms on its queue once the generator has finished: it takes what is queued and returns. Without this, every remote Doris query would wait 100 ms before the first DoGet of each backend holding an endpoint (measured on a one-row lookup, 30 runs right after an FE restart: median 160 ms without it, 58 ms with it); hive / iceberg batch scans lose the same 100 ms tail per backend.[fix](be)A backend no longer takes a cached thrift connection to a server that closed it. Once a frontend restarts, every connection a backend had cached to it is dead. The split fetch (fetchSplitBatch) took one from the cache and failed the scan: unlike the other backend-to-frontend calls it does not reopen and send again, and it cannot, because the frontend hands out the splits of a batch as it answers, so a request sent again after a lost answer would skip a batch. The client cache now looks at a cached client before it hands it out - a non-blocking peek at its socket: the server's end of stream, or an error, is waiting - and closes it instead of handing it out, so every caller gets a live connection. Without this, every remote Doris query, whose backends all fetch after this PR, would fail right after its frontend restarts, as batch-mode hive / iceberg scans already could.[fix](remote-doris)The remote Doris scan becomes a batch split-source scan withneedsSampleSplit() == false.convertPredicate; EXPLAIN shows it unchanged) and one scan range per backend pointing at a split source.startSplit()runs at dispatch. It tries the remote FEs in turn as before, stopping early if the scan was stopped meanwhile. It hands the session to the split assignment and queues the endpoints as splits.query_timeout_sec, as GetFlightInfo has. The coordinator runs the remote query while it holds the query's admission slot, and the handshake used to wait without one: a remote frontend that accepts the connection but never answers would hold the slot, and the connection thread, for good - neither KILL nor the query's timeout interrupts the wait. A dispatch's remote work is now bounded byquery_retry_countx (connect + handshake + GetFlightInfo + closing a failed attempt): about 3 x 70 s with the catalog defaults.isExplainStatement(); the scan's session fields and itsstop(),coordinatorMustOutliveDispatch()andcannotBeRedispatched()overrides; theStatementContextfallback (stopScanNodeAtClose,handOverScanNodesToDeferredCoordinator, the step inclose()); and the hand-over in the Arrow Flight deferral gate. The split assignment is what keeps a deferred Flight coordinator alive, as for any batch scan.[fix](fe)A batch scan that plans with its first split (hive / iceberg / paimon / MaxCompute) starts its split generation while planning, and nothing stopped it when the plan was dropped: every successfulCREATE JOB ... DO INSERT ... SELECT, every task of a streaming insert job whose SELECT reads such a table besides the TVF (before()plans the statement only to rewrite the TVF), an INSERT whose label is taken, a statement refused by a SQL block rule after planning, an http_stream load refused after planning, the plan an INSERT drops when its target table changes, a DELETE whose predicate reads another table (it falls back toDeleteFromUsingCommand, which plans again). The streaming generator of an iceberg scan then waits forever on a full backend queue, holding one of the 64 sharedscheduleExecutorthreads until the FE restarts; a streaming job did that everymax_interval(10 s).exec()takes it over (SplitAssignment.start).ScanNode.stopAllUndispatched, which logs a failed split generation as one instead of throwing it as a failure to stop): EXPLAIN, the INSERT OVERWRITE probe, the INSERT planned again, DELETE, the plan a streaming task'sbefore()builds to rewrite the TVF.StatementContext.stopUndispatchedSplitAssignments):StatementContext.close(); the end of a binaryCOM_STMT_EXECUTE, whose context the prepared statement keeps until its next execution, so it must not keep the plan alive through an assignment; an http_stream request; every attempt of a streaming insert task (on the task thread, a canceled attempt too) and the attempt a cloud re-plan retries.[fix](fe)A backend's fetch created an unbounded split queue when it came before the generator's first split for that backend, so the generator never waited for that backend. Both now create the same bounded queue. Found while testing item 6.[fix](regression-framework)The connections of the thread-local Doris accessors are closed once their thread ends (OpenedDorisConnections, swept from all three accessors, with strays kept until their thread ends and closed at the end of the run), and Awaitility no longer blames the suite that happens to be awaiting for another thread's uncaught exception.[fix](regression-framework)A suite fails for an uncaught exception on a thread it started (UncaughtThreadFailures), including a thread that a task ofSuite.thread()starts, which runs as the suite that submitted it; a suite that fails on something else carries those failures along as suppressed exceptions.[fix](fe)Unrelated to remote Doris, found when an External run of this PR lost its FE: dropping or altering an ADBC catalog while a statement or a column-statistics load was still inside its driver crashed the FE (SIGSEGV insqlite3FindTable).AdbcClient.close()closed the nativeAdbcDatabaseat once, and the JNI driver first closes every connection still open on it, from the closing thread, under the threads using them. Nowclose()only refuses new calls, and the native driver is released by whoever leaves last:close()itself when no call is in flight, otherwise the last call to leave, which logs a failed release instead of failing its own work.close()still does not wait, so DROP / ALTER CATALOG are not held up. Self-contained: it can be picked without the rest of this PR.No new outbound request: the existing Flight SQL request to the catalog's
fe_arrow_hostsnow happens only when the local query runs, and fewer times than before.What changes besides:
Failed to execute querymessage.Results
Remote queries each statement ran on the remote FE, counted from the remote FE's audit log, plus the catalog user's Flight SQL sessions left open after each statement. Local cluster, one FE and one BE, with the catalog pointing back at itself. "Before" is the previous head of this PR, which behaves like master on these statements.
SELECT ... FROM remoteEXPLAIN SELECT .../* a comment first */ EXPLAIN SELECT ...INSERT OVERWRITE TABLE t SELECT ... FROM remoteINSERT ... WITH LABEL l SELECT ...(label free)CREATE JOB ... STARTS '2099-01-01' DO INSERT ... SELECT FROM remoteCREATE JOB ... AT CURRENT_TIMESTAMP(creation + its task)NullPointerException ... "originStmt"An Arrow Flight SQL client reading a remote Doris table: 3 rows returned, the remote session open (1) while the deferred query waits, and 0 after the client's next request finalizes it.
Classes, and how they call each other
SplitAssignment:registerSource(planning),start(the coordinator takes it over; registers sources, theninit->SplitGenerator.startSplit),startWhilePlanning(a scan that plans with its first split),addCloseable,stop(unregisters, drops the queued splits, closes, rethrows a generation failure last),stopIfNotDispatched(the statement's end: the same unless a coordinator took it over, logging a generation failure instead of throwing it).ScanNode:start/stopdelegate to the split assignment;startAll/stopAllfor the coordinators;stopUndispatched/stopAllUndispatched(SplitAssignment.stopIfNotDispatched, no rethrow) for a plan the statement drops;coordinatorMustOutliveDispatch(unchanged:splitAssignment != null);cannotBeRedispatched(splitAssignment != null && splitAssignment.isStop()).FileQueryScanNode.createScanRangeLocations: batch mode creates the assignment and one split source per backend; it starts the assignment only whenneedsSampleSplit(), registering it with the statement first.Coordinator.exec/NereidsCoordinator.exec:ScanNode.startAllafter admission, before the fragments are built;close/cancel:ScanNode.stopAll.RemoteDorisScanNode:convertPredicatebuilds the query;isBatchModetrue,needsSampleSplitfalse,numApproximateSplits0;startSplit->executeQuery->executeFlightSqlQuery(RemoteDorisFlightSession.open+execute,splitAssignment.addCloseable(session)) ->addToQueue(endpoints)->finishSchedule.RemoteDorisFlightSession:openbounds the handshake byquery_timeout_sec(it had no deadline); CloseSession bounded to 5s onclose(), as before.StatementContext: no longer knows about scan nodes; it keeps the split assignments its planning started, andstopUndispatchedSplitAssignmentsstops those no coordinator took over and forgets all of them. Called byclose(),MysqlConnectProcessor.handleExecute(end of a binaryCOM_STMT_EXECUTE),FrontendServiceImpl.initHttpStreamPlan,StmtExecutor.queryRetry(before a re-plan) andStreamingInsertTask.endAttempt(fromAbstractStreamingTask.execute's finally).StmtExecutor.planCannotBeRedispatchedreads the general rule.ScanNode.stopAllUndispatchedon a plan dropped mid-statement:ExplainCommand,InsertOverwriteTableCommand(probe),InsertIntoTableCommand.initPlan(re-plan),DeleteFromCommand.run,StreamingInsertTask.before.ClientCacheHelper::get_client(BE, item 4): a cached client whose server closed the connection (ThriftClientImpl::peer_closed) is closed, and the next cached one tried or a new one opened;reopen_clientshares the closing (_close_client).AdbcClient(item 10):withConnection->enter(counts the call in, opens the database on first use) ...leave(counts it out; the last one out of a closed client callsrelease);closerefuses new calls and callsreleaseonly when no call is in flight.Release note
A remote Doris catalog (
use_arrow_flight = true) runs the query on the remote cluster only when the local query runs. EXPLAIN, statements that fail before they run, and the plans INSERT OVERWRITE and CREATE JOB only inspect no longer reach the remote cluster, and INSERT OVERWRITE runs the remote query once instead of twice. Jobs and streaming insert jobs can read remote Doris tables; before, they failed with a NullPointerException. A query with a batch-mode external table scan that fails with an RPC error is no longer retried with the same plan, which could only fail with "Split source ... is released"; it reports the original error. A CREATE JOB, a streaming insert job, or a statement refused or failing after planning, whose plan reads a large iceberg table in batch mode no longer leaves a frontend thread generating splits forever; enough of them stalled every batch-mode external table scan on that frontend. The frontend no longer keeps most of the splits of a huge batch-mode scan in memory when a backend fetches before it was assigned a split. A query that reads remote Doris, or a batch-mode external table, no longer fails with "Failed to get batch of split source" right after the frontend running it restarts. A remote Doris query whose remote frontend accepts the connection but never answers gives up after the catalog'squery_timeout_secinstead of waiting for good. Fix an FE crash (SIGSEGV in the ADBC driver) when an ADBC catalog is dropped or altered while a statement or a statistics load is still reading it.Check List (For Author)
Test
Unit tests, fe-core:
SplitAssignmentTest,StatementContextTest,MysqlConnectProcessorExecuteEndTest,MysqlConnectProcessorCursorFetchTest,StreamingInsertTaskAttemptEndTest,StreamingInsertTaskAuditTest,DeleteFromCommandTest,FrontendServiceImplHttpStreamPlanTest,StmtExecutorReplanRetryTest,ScanNodeDispatchTest,RemoteDorisScanNodeTest,ConnectorStatementScopeTest,ArrowFlightDeferralGateTest,OldCoordinatorTest,NereidsCoordinatorTest,ExecuteCommandTest,StmtExecutorTest,ProtocolCapabilityWiringTestand thePluginDrivenScanNodebatch mode / scan profile / read txn / compatibility tests: 179, all passing; checkstyle 0.AdbcClientTest: 11, all passing with the real JNI bridge and SQLite driver (0 skipped). Regression framework: 58 (those two commits are unchanged since that run), includingSuiteContextDorisConnectionsTestandSuiteThreadFailuresTest, which run the accessors and whole suites. BE:client_cache_test, 3 cases, all passing under ASAN; with the check inget_clienttaken out, the stale-connection case fails. Each fix was checked against its test: undoing it fails the test, for every fix but one - the plan an INSERT drops when its target table changes between planning and locking has no unit test. Regression, on a local cluster: the 13 suites ofexternal_table_p0/remote_doris, includingtest_remote_doris_flight_session, which binds a SQL block rule that refuses every remote query of the catalog user and checks that EXPLAIN (also behind a comment), a refused INSERT and CREATE JOB never reach the remote FE, and that a job's insert task reads the remote table, its rows going through a generated.out;prepared_stmt_p0(server-side prepared statements, through the new end of a binaryCOM_STMT_EXECUTE),load_p0/http_stream, anddelete_p0/test_delete_where_in(a DELETE falling back toDeleteFromUsingCommand): 29 suites, all passing. Manual, on the same cluster: the per-statement counts and the Arrow Flight client check above; 20 remote Doris queries, then three FE restarts with 20 remote Doris queries right after each: none failed, and the backend closed 9 cached connections the restarted frontend had closed (item 4) - earlier local runs without item 4 logged remote Doris queries failing right after a frontend restart withFailed to get batch of split source; a remote Doris catalog whosefe_arrow_hostsaccepts the connection and never answers (query_timeout_sec3,query_retry_count2): the query fails after 6 s withFailed to execute queryinstead of waiting for good (item 5). The batch-mode hive / iceberg / paimon suites were not run locally (no docker environment for them): their planning and split generation are unchanged, but who stops that generation is not (item 6), so those suites are the ones to watch in CI.Behavior changed:
Does this need documentation?
Check List (For Reviewer who merge this PR)