Expected behavior
When a data file that was already opened and cached by ZipDataCacheProvider becomes unreadable mid-run (the inode is replaced or unlinked on the mount while the process holds the handle, so the next read fails with IOException: Stale file handle), the engine should either re-open the file on the next request or fail the run. A data file that exists and is valid on disk should not be served as "no data" for the rest of the backtest.
Actual behavior
Once a read on a cached CachedZipFile throws, ZipDataCacheProvider.Fetch() logs Corrupt zip file/entry and returns null, but leaves the poisoned CachedZipFile in _zipFileCache. Every later Fetch() for that file goes back to the same instance, calls existingZip.Refresh() first (which resets its cache timestamp, so CleanCache() never reclaims it), throws again, and returns null again. The same file keeps failing until the run ends; in one run we counted 432 failures on a single qqq.zip over 40 minutes, the last one seconds before Analysis Completed.
A null from Fetch() makes TextSubscriptionDataSourceReader.Read() yield break (no bars, no exception), so a History() call for a symbol whose daily file is in that state returns an empty DataFrame, OnData sees nothing wrong, and the backtest reports Completed with results computed from missing history. Nothing distinguishes this from a symbol that genuinely has no data, so the algorithm cannot detect it.
The file on disk is fine the whole time: a fresh process opens it without error, and every backtest launched after the swap reads it cleanly. Only the process holding the pre-swap handle is affected, and it never recovers.
Reproduce
Observed in cloud backtests on 2026-09-28 (LEAN 2.5.0.0.18130), Intercom conversation 215476143715485; six long backtests in one project reading equity/usa/daily/*.zip were hit at three wall-clock instants during the afternoon (each coinciding with the data folder being replaced under the running nodes). Syslog of one of them, first hit on the SPY daily file:
2026-09-28T15:04:20 ERROR:: ZipDataCacheProvider.Fetch(): Corrupt zip file/entry: /Data/equity/usa/daily/spy.zip# Error: Ionic.Zip.ZipException: Cannot read that as a ZipFile
---> System.IO.IOException: Stale file handle : '/cache/equity/usa/daily/spy.zip'
at System.IO.Strategies.OSFileStreamStrategy.Read(Span`1 buffer)
at Ionic.Zip.SharedUtilities._ReadFourBytes(Stream s, String message)
at Ionic.Zip.ZipFile.ReadIntoInstance(ZipFile zf)
--- End of inner exception stack trace ---
at QuantConnect.Lean.Engine.DataFeeds.ZipDataCacheProvider.CreateEntryStream(CachedZipFile zipFile, String entryName, String fileName)
at QuantConnect.Lean.Engine.DataFeeds.ZipDataCacheProvider.Fetch(String key)
2026-09-28T15:04:51 ERROR:: ZipDataCacheProvider.Fetch(): Corrupt zip file/entry: /Data/equity/usa/daily/spy.zip# Error: Ionic.Zip.ZipException: Cannot read that as a ZipFile
---> System.IO.IOException: Stale file handle : '/cache/equity/usa/daily/spy.zip'
The run then logged the same error for spy.zip on every following simulated day, and the algorithm's own log shows its History([SPY, QQQ], 300, Resolution.Daily) returning an empty DataFrame from that simulated date to the end of the run, with no exception. Across the six runs the same file failed between 3 and 3,137 times each, always starting at the wall-clock instant of the swap and never recovering.
Reproduction outside the cloud, on any Linux host with the data folder on NFS (or any mount that gives ESTALE on an unlinked open file): start a long daily-resolution backtest that calls History() every day, and while it runs replace the equity/usa/daily directory by rename (mv daily daily-old && mv daily-new daily && rm -rf daily-old, i.e. new inodes). On a local ext4 disk the unlinked inode stays readable so the symptom does not appear; a unit test can simulate it with a Stream that starts throwing IOException after the CachedZipFile has been constructed.
Evidence in the code (master at the time of filing)
Engine/DataFeeds/ZipDataCacheProvider.cs, Fetch(): the catch for a cached zip (lines 96-107) logs ZipException/ZlibException and falls through to return stream (null); any other exception reaches the outer catch (lines 111-121) which logs Inner try/catch and returns null. Neither path removes existingZip from _zipFileCache or disposes it.
- Same method, line 93:
existingZip.Refresh() runs before every CreateEntryStream, so a failing entry has its _dateCached bumped on each attempt and CleanCache() (lines 229-260, zip.Value.Uncache(clearCacheIfOlderThan)) never evicts it while the algorithm keeps asking for that symbol.
Engine/DataFeeds/TextSubscriptionDataSourceReader.cs, lines 100-103: a null reader is yield break, so the failure is indistinguishable from "no file" for SubscriptionDataReader, HistoryProvider and the algorithm.
Proposed change
- In
Fetch() (and CacheAndCreateEntryStream), when creating the entry stream throws for a cached instance, dispose it and _zipFileCache.TryRemove(filename, ...) before returning, so the next request re-opens the file through IDataProvider (which, after a swap, succeeds: the new file is valid). This is the same recovery the code already does for a Modified entry (CreateEntryStream, lines 329-336).
- Distinguish an I/O failure from a missing file: an
IOException reading a file the provider already resolved should not be reported to the reader as an empty result. Either rethrow it (a failed backtest is recoverable, a silently wrong one is not) or surface it through the data-monitor as a failed data request so it appears in the run's DATA USAGE summary and the error field.
Open questions
- Whether the cloud build's read path (which adds a fast-read step before the managed
ZipFile fallback) needs the same eviction in both places; the public method structure is the same.
- Whether the daily/hour "one rolling zip per ticker" layout should keep the file handle at all between requests, given the file can legitimately be rewritten while a run is in progress.
Expected behavior
When a data file that was already opened and cached by
ZipDataCacheProviderbecomes unreadable mid-run (the inode is replaced or unlinked on the mount while the process holds the handle, so the next read fails withIOException: Stale file handle), the engine should either re-open the file on the next request or fail the run. A data file that exists and is valid on disk should not be served as "no data" for the rest of the backtest.Actual behavior
Once a read on a cached
CachedZipFilethrows,ZipDataCacheProvider.Fetch()logsCorrupt zip file/entryand returnsnull, but leaves the poisonedCachedZipFilein_zipFileCache. Every laterFetch()for that file goes back to the same instance, callsexistingZip.Refresh()first (which resets its cache timestamp, soCleanCache()never reclaims it), throws again, and returnsnullagain. The same file keeps failing until the run ends; in one run we counted 432 failures on a singleqqq.zipover 40 minutes, the last one seconds beforeAnalysis Completed.A
nullfromFetch()makesTextSubscriptionDataSourceReader.Read()yield break(no bars, no exception), so aHistory()call for a symbol whose daily file is in that state returns an empty DataFrame,OnDatasees nothing wrong, and the backtest reportsCompletedwith results computed from missing history. Nothing distinguishes this from a symbol that genuinely has no data, so the algorithm cannot detect it.The file on disk is fine the whole time: a fresh process opens it without error, and every backtest launched after the swap reads it cleanly. Only the process holding the pre-swap handle is affected, and it never recovers.
Reproduce
Observed in cloud backtests on 2026-09-28 (LEAN 2.5.0.0.18130), Intercom conversation 215476143715485; six long backtests in one project reading
equity/usa/daily/*.zipwere hit at three wall-clock instants during the afternoon (each coinciding with the data folder being replaced under the running nodes). Syslog of one of them, first hit on the SPY daily file:The run then logged the same error for
spy.zipon every following simulated day, and the algorithm's own log shows itsHistory([SPY, QQQ], 300, Resolution.Daily)returning an empty DataFrame from that simulated date to the end of the run, with no exception. Across the six runs the same file failed between 3 and 3,137 times each, always starting at the wall-clock instant of the swap and never recovering.Reproduction outside the cloud, on any Linux host with the data folder on NFS (or any mount that gives ESTALE on an unlinked open file): start a long daily-resolution backtest that calls
History()every day, and while it runs replace theequity/usa/dailydirectory by rename (mv daily daily-old && mv daily-new daily && rm -rf daily-old, i.e. new inodes). On a local ext4 disk the unlinked inode stays readable so the symptom does not appear; a unit test can simulate it with aStreamthat starts throwingIOExceptionafter theCachedZipFilehas been constructed.Evidence in the code (master at the time of filing)
Engine/DataFeeds/ZipDataCacheProvider.cs,Fetch(): thecatchfor a cached zip (lines 96-107) logsZipException/ZlibExceptionand falls through toreturn stream(null); any other exception reaches the outer catch (lines 111-121) which logsInner try/catchand returnsnull. Neither path removesexistingZipfrom_zipFileCacheor disposes it.existingZip.Refresh()runs before everyCreateEntryStream, so a failing entry has its_dateCachedbumped on each attempt andCleanCache()(lines 229-260,zip.Value.Uncache(clearCacheIfOlderThan)) never evicts it while the algorithm keeps asking for that symbol.Engine/DataFeeds/TextSubscriptionDataSourceReader.cs, lines 100-103: anullreader isyield break, so the failure is indistinguishable from "no file" forSubscriptionDataReader,HistoryProviderand the algorithm.Proposed change
Fetch()(andCacheAndCreateEntryStream), when creating the entry stream throws for a cached instance, dispose it and_zipFileCache.TryRemove(filename, ...)before returning, so the next request re-opens the file throughIDataProvider(which, after a swap, succeeds: the new file is valid). This is the same recovery the code already does for aModifiedentry (CreateEntryStream, lines 329-336).IOExceptionreading a file the provider already resolved should not be reported to the reader as an empty result. Either rethrow it (a failed backtest is recoverable, a silently wrong one is not) or surface it through the data-monitor as a failed data request so it appears in the run'sDATA USAGEsummary and theerrorfield.Open questions
ZipFilefallback) needs the same eviction in both places; the public method structure is the same.