Retrieval-augmented Large Language Models for Financial Time Series Forecasting
This repository contains the code for FinSeer, a domain-specific retriever for financial time-series forecasting, used in a retrieval-augmented generation (RAG) framework for stock movement prediction. FinSeer is trained with candidate selection refined by LLM feedback and a similarity-driven objective, and the retrieved sequences are fed into StockLLM, an LLM fine-tuned for stock movement prediction. The code covers indicator calculation, retriever training, and stock movement prediction with and without retrieval.
| Resource | Type | Description |
|---|---|---|
| FinSeer collection | Collection | All FinSeer resources |
| TheFinAI/FinSeer | Model | Our retriever |
| TheFinAI/StockLLM | Model | Our fine-tuned stock LLM |
src/
├── 1_calulate_indicators/ # financial indicator calculation
├── 2_train_retriever/ # LLM feedback scoring and candidate selection
└── 3_stock_movement_prediction/ # embeddings, similarity, and RAG prediction
# for baseline RAG models and retriever training
pip install InstructorEmbedding
pip install -U FlagEmbedding
pip install sentence-transformers==2.2.2
pip install protobuf==3.20.0
pip install yahoo-finance
python -m pip install -U angle-emb
pip install transformers==4.33.2 # UAEStep 1. Get LLM feedback scores (src/2_train_retriever/get_llm_feedback_scores.py)
dataset:acl18,bigdata22,stock23target: the file to save, a JSON with LLM probability scores
parser = argparse.ArgumentParser(description='test')
parser.add_argument('--dataset', default='acl18', type=str)
parser.add_argument('--target', default='acl18.scored.json', type=str)
args = parser.parse_args()
get_all_scores(llm='StockLLM') # llama family are all supported for this codeStep 2. Select positive and negative candidates (src/2_train_retriever/select_positive_and_combine.py)
Before this step, you should have generated acl18.scored.json, bigdata22.scored.json and stock23.scored.json.
This step selects candidates for all three datasets and generates a combined train.scored.json file.
Then, follow the steps in this link to fine-tune your own FinSeer using the train.scored.json data.
Step 1. Get embeddings of queries and candidates (src/3_stock_movement_prediction/get_embeddings.py)
q_or_c: query or candidate; we generate the embeddings of query sequences and candidate sequences separately
parser = argparse.ArgumentParser(description='test')
parser.add_argument('--test_dataset', default='bigdata22', type=str)
parser.add_argument('--embedding_model', default='e5',
choices=['instructor', 'uae', 'bge', 'llm_embedder', 'e5', 'FinSeer'])
parser.add_argument('--q_or_c', default='candidate')
args = parser.parse_args()Step 2. Calculate the similarity of queries and qualified candidates (and get the top-5 related candidates).
Step 3. Predict stock movement with or without retrieval. This is what our three files do:
1_no_retrieval.py2_random_retrieval.py3_similarity_retrieval.py
If you find FinSeer useful, please cite:
@misc{xiao2025retrievalaugmented,
title={Retrieval-augmented Large Language Models for Financial Time Series Forecasting},
author={Mengxi Xiao and Zihao Jiang and Lingfei Qian and Zhengyu Chen and Yueru He and Yijing Xu and Yuecheng Jiang and Dong Li and Ruey-Ling Weng and Min Peng and Jimin Huang and Sophia Ananiadou and Qianqian Xie},
year={2025},
eprint={2502.05878},
archivePrefix={arXiv},
primaryClass={cs.CL}
}The code in this repository is released under the MIT License. Datasets and models on Hugging Face keep their own licenses, stated on each card.
Built by The Fin AI · Hugging Face · GitHub