منابع کتاب Foundations of Large Language Models — بخش اول

منابع کتاب Foundations of Large Language Models — بخش اول

توسط admin | گروه هوش مصنوعی | 1405/05/18

نظرات 0

منابع کتاب Foundations of Large Language Models — بخش اول

عنوان اصلی کتاب: Foundations of Large Language Models

نویسندگان: Tong Xiao و Jingbo Zhu

سازمان: NLP Lab, Northeastern University & NiuTrans Research

زبان اصلی: انگلیسی

بازهٔ منبع: صفحات PDF 210 تا 220

مجوز منبع: Creative Commons Attribution-NonCommercial 4.0 (CC BY-NC 4.0)

تاریخ ترجمه: 1405/05/18 / 2026-08-09

اعتبار ترجمه: ترجمه با کمک هوش مصنوعی

منابع کتاب — بخش اول

اطلاعات کتاب‌شناختی برای حفظ قابلیت ارجاع، با نام نویسندگان، عنوان آثار، محل انتشار، شماره صفحات و سال انتشارِ منبع اصلی منتقل شده است. عنوان‌ها و شناسه‌های آثار در شکل اصلی خود حفظ شده‌اند.

[Ainslie et al., 2020] Joshua Ainslie, Santiago Ontanon, Chris Alberti, Vaclav Cvicek, Zachary Fisher,
  Philip Pham, Anirudh Ravula, Sumit Sanghai, Qifan Wang, and Li Yang. Etc: Encoding long and
  structured inputs in transformers. In Proceedings of the 2020 Conference on Empirical Methods in
  Natural Language Processing (EMNLP), pages 268–284, 2020.
[Ainslie et al., 2023] Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico
  Lebron, and Sumit Sanghai. Gqa: Training generalized multi-query transformer models from multi-
  head checkpoints. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language
  Processing, pages 4895–4901, 2023.
[Akyürek et al., 2023] Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou.
  What learning algorithm is in-context learning? investigations with linear models. In Proceedings of
  The Eleventh International Conference on Learning Representations, 2023.
[Alabdulmohsin et al., 2022] Ibrahim M Alabdulmohsin, Behnam Neyshabur, and Xiaohua Zhai. Revisit-
  ing neural scaling laws in language and vision. Advances in Neural Information Processing Systems, 35:
  22300–22312, 2022.
[Allal et al., 2024] Loubna Ben Allal, Anton Lozhkov, and Daniel van Strien. cosmopedia: how to create
  large-scale synthetic data for pre-training. https://huggingface.co/blog/cosmopedia, 2024.
[Almazrouei et al., 2023] Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cap-
  pelli, Ruxandra Cojocaru, Mérouane Debbah, Étienne Goffinet, Daniel Hesslow, Julien Launay, Quentin
  Malartic, Daniele Mazzotta, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo. The falcon
  series of open language models. arXiv preprint arXiv:2311.16867, 2023.
[Andreas et al., 2016] Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein. Neural module
  networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages
  39–48, 2016.
[Arjovsky et al., 2016] Martin Arjovsky, Amar Shah, and Yoshua Bengio. Unitary evolution recurrent
  neural networks. In International conference on machine learning, pages 1120–1128, 2016.
[Aschenbrenner, 2024] Leopold Aschenbrenner. Situational awareness: The decade ahead, 2024. URL
  https://situational-awareness.ai/.
[Askell et al., 2021] Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan,
  Andy Jones, Nicholas Joseph, Benjamin Mann, Nova DasSarma, Nelson Elhage, Zac Hatfield-Dodds,
  Danny Hernandez, Jackson Kernion, Kamal Ndousse, Catherine Olsson, Dario Amodei, Tom B. Brown,
  Jack Clark, Sam McCandlish, Chris Olah, and Jared Kaplan. A general language assistant as a laboratory
  for alignment. arXiv preprint arXiv:2112.00861, 2021.
[Bach et al., 2022] Stephen H. Bach, Victor Sanh, Zheng Xin Yong, Albert Webson, Colin Raffel, Nihal V.
  Nayak, Abheesht Sharma, Taewoon Kim, M. Saiful Bari, Thibault Févry, Zaid Alyafeai, Manan Dey,
  Andrea Santilli, Zhiqing Sun, Srulik Ben-David, Canwen Xu, Gunjan Chhablani, Han Wang, Jason Alan
  Fries, Maged Saeed AlShaibani, Shanya Sharma, Urmish Thakker, Khalid Almubarak, Xiangru Tang,
  Dragomir R. Radev, Mike Tian-Jian Jiang, and Alexander M. Rush. Promptsource: An integrated de-
  velopment environment and repository for natural language prompts. In Proceedings of the 60th Annual
  Meeting of the Association for Computational Linguistics: System Demonstrations, pages 93–104, 2022.
[Bengio et al., 2003] Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Jauvin. A neural
  probabilistic language model. Journal of Machine Learning Research, 3:1137–1155, 2003.
[Bengio et al., 2006] Yoshua Bengio, Pascal Lamblin, Dan Popovici, and Hugo Larochelle. Greedy layer-
  wise training of deep networks. Advances in neural information processing systems, 19, 2006.
[Bengio et al., 2024] Yoshua Bengio, Geoffrey Hinton, Andrew Yao, Dawn Song, Pieter Abbeel, Trevor
  Darrell, Yuval Noah Harari, Ya-Qin Zhang, Lan Xue, Shai Shalev-Shwartz, Gillian K. Hadfield, Jeff

  Clune, Tegan Maharaj, Frank Hutter, Atilim Gunes Baydin, Sheila A. McIlraith, Qiqi Gao, Ashwin
  Acharya, David Krueger, Anca Dragan, Philip Torr, Stuart Russell, Daniel Kahneman, Jan Markus
  Brauner, and Sören Mindermann. Managing extreme ai risks amid rapid progress. Science, 384(6698):
  842–845, 2024.
[Bentivogli and Giampiccolo, 2011] Luisa Bentivogli and Danilo Giampiccolo. Pascal recognizing textual
  entailment challenge (rte-7) at tac 2011. https://tac.nist.gov/2011/RTE/, 2011.
[Besta et al., 2024] Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski,
  Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, and Torsten
  Hoefler. Graph of thoughts: Solving elaborate problems with large language models. In Proceedings of
  the AAAI Conference on Artificial Intelligence, volume 38, pages 17682–17690, 2024.
[Biderman et al., 2021] Stella Biderman, Sid Black, Charles Foster, Leo Gao, Eric Hallahan, Horace He,
  Ben Wang, and Phil Wang. Rotary embeddings: A relative revolution. https://blog.eleuther.ai/
  rotary-embeddings/, 2021.
[Bishop, 2006] Christopher M. Bishop. Pattern Recognition and Machine Learning. Springer, 2006.
[Blum and Mitchell, 1998] Avrim Blum and Tom Mitchell. Combining labeled and unlabeled data with
  co-training. In Proceedings of the eleventh annual conference on Computational learning theory, pages
  92–100, 1998.
[Bradley and Terry, 1952] Ralph Allan Bradley and Milton E. Terry. Rank analysis of incomplete block
  designs: I. the method of paired comparisons. Biometrika, 39(3/4):324–345, 1952.
[Brandon et al., 2024] William Brandon, Mayank Mishra, Aniruddha Nrusimha, Rameswar Panda, and
  Jonathan Ragan Kelly. Reducing transformer key-value cache size with cross-layer attention. arXiv
  preprint arXiv:2405.12981, 2024.
[Brill, 1992] Eric Brill. A simple rule-based part of speech tagger. In Speech and Natural Language:
  Proceedings of a Workshop Held at Harriman, New York, February 23-26, 1992, 1992.
[Brown et al., 1993] Peter F. Brown, Stephen A. Della Pietra, Vincent J. Della Pietra, and Robert L. Mercer.
  The mathematics of statistical machine translation: Parameter estimation. Computational Linguistics,
  19(2):263–311, 1993.
[Brown et al., 2020] Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla
  Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel
  Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey
  Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess,
  Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei.
  Language models are few-shot learners. Advances in neural information processing systems, 33:1877–
  1901, 2020.
[Bubeck et al., 2023] Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric
  Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott M. Lundberg, Harsha Nori, Hamid
  Palangi, Marco Túlio Ribeiro, and Yi Zhang. Sparks of artificial general intelligence: Early experiments
  with gpt-4. arXiv preprint arXiv:2303.12712, 2023.
[Bulatov et al., 2022] Aydar Bulatov, Yury Kuratov, and Mikhail Burtsev. Recurrent memory transformer.
  Advances in Neural Information Processing Systems, 35:11079–11091, 2022.
[Burges et al., 2005] Chris Burges, Tal Shaked, Erin Renshaw, Ari Lazier, Matt Deeds, Nicole Hamilton,
  and Greg Hullender. Learning to rank using gradient descent. In Proceedings of the 22nd international
  conference on Machine learning, pages 89–96, 2005.
[Burns et al., 2023] Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold
  Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, Ilya Sutskever, and Jeff Wu.
  Weak-to-strong generalization: Eliciting strong capabilities with weak supervision. arXiv preprint
  arXiv:2312.09390, 2023a.

[Burns et al., 2023] Collin Burns, Jan Leike, Leopold Aschenbrenner, Jeffrey Wu, Pavel Izmailov, Leo
  Gao, Bowen Baker, and Jan Hendrik Kirchner. Weak-to-strong generalization, 2023b. URL https://
  https://openai.com/index/weak-to-strong-generalization.
[Caballero et al., 2023] Ethan Caballero, Kshitij Gupta, Irina Rish, and David Krueger. Broken neural
  scaling laws. In ICLR 2023 Workshop on Mathematical and Empirical Understanding of Foundation
  Models, 2023.
[Cao et al., 2007] Zhe Cao, Tao Qin, Tie-Yan Liu, Ming-Feng Tsai, and Hang Li. Learning to rank: from
  pairwise approach to listwise approach. In Proceedings of the 24th international conference on Machine
  learning, pages 129–136, 2007.
[Chang et al., 2024] Kaiyan Chang, Songcheng Xu, Chenglong Wang, Yingfeng Luo, Tong Xiao, and
  Jingbo Zhu. Efficient prompting methods for large language models: A survey. arXiv preprint
  arXiv:2404.01077, 2024.
[Charniak, 1997] Eugene Charniak. Statistical parsing with a context-free grammar and word statistics.
  AAAI/IAAI, 2005(598-603):18, 1997.
[Chen et al., 2023] Banghao Chen, Zhaofeng Zhang, Nicolas Langrené, and Shengxin Zhu. Unleashing
  the potential of prompt engineering in large language models: a comprehensive review. arXiv preprint
  arXiv:2310.14735, 2023a.
[Chen et al., 2023] Lichang Chen, Shiyang Li, Jun Yan, Hai Wang, Kalpa Gunaratna, Vikas Yadav, Zheng
  Tang, Vijay Srinivasan, Tianyi Zhou, Heng Huang, and Hongxia Jin. Alpagasus: Training a better alpaca
  with fewer data. arXiv preprint arXiv:2307.08701, 2023b.
[Chen et al., 2024] Lichang Chen, Shiyang Li, Jun Yan, Hai Wang, Kalpa Gunaratna, Vikas Yadav, Zheng
  Tang, Vijay Srinivasan, Tianyi Zhou, Heng Huang, and Hongxia Jin. Alpagasus: Training a better alpaca
  with fewer data. In The Twelfth International Conference on Learning Representations, 2024a.
[Chen et al., 2023] Shouyuan Chen, Sherman Wong, Liangjian Chen, and Yuandong Tian. Extending
  context window of large language models via positional interpolation. arXiv preprint arXiv:2306.15595,
  2023c.
[Chen et al., 2020] Tianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu, Yang Zhang, Zhangyang
  Wang, and Michael Carbin. The lottery ticket hypothesis for pre-trained bert networks. Advances in
  neural information processing systems, 33:15834–15846, 2020.
[Chen et al., 2024] Zixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji, and Quanquan Gu. Self-play fine-
  tuning converts weak language models to strong language models. arXiv preprint arXiv:2401.01335,
  2024b.
[Chevalier et al., 2023] Alexis Chevalier, Alexander Wettig, Anirudh Ajith, and Danqi Chen. Adapting
  language models to compress contexts. In Proceedings of the 2023 Conference on Empirical Methods
  in Natural Language Processing, pages 3829–3846, 2023.
[Chi et al., 2022] Ta-Chung Chi, Ting-Han Fan, Peter J Ramadge, and Alexander Rudnicky. Kerple:
  Kernelized relative positional embedding for length extrapolation. Advances in Neural Information Pro-
  cessing Systems, 35:8386–8399, 2022.
[Chi et al., 2023] Ta-Chung Chi, Ting-Han Fan, Alexander Rudnicky, and Peter Ramadge. Dissecting
  transformer length extrapolation via the lens of receptive field analysis. In Proceedings of the 61st Annual
  Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 13522–13537,
  2023.
[Chiang et al., 2023] Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lian-
  min Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. Vicuna:
  An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023. URL https://
  lmsys.org/blog/2023-03-30-vicuna/.
[Chowdhery et al., 2022] Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav

  Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker
  Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam
  Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury,
  Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghe-
  mawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus,
  Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan
  Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankara-
  narayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov,
  Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta,
  Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. Palm: Scaling
  language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022.
[Christiano et al., 2017] Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario
  Amodei. Deep reinforcement learning from human preferences. Advances in neural information pro-
  cessing systems, 30, 2017.
[Chu et al., 2023] Zheng Chu, Jingchang Chen, Qianglong Chen, Weijiang Yu, Tao He, Haotian Wang,
  Weihua Peng, Ming Liu, Bing Qin, and Ting Liu. A survey of chain of thought reasoning: Advances,
  frontiers and future. arXiv preprint arXiv:2309.15402, 2023.
[Chung et al., 2022] Hyung Won Chung, Le Hou, S. Longpre, Barret Zoph, Yi Tay, William Fedus,
  Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu,
  Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Dasha Valter, Sharan Narang, Gau-
  rav Mishra, Adams Wei Yu, Vincent Zhao, Yanping Huang, Andrew M. Dai, Hongkun Yu, Slav Petrov,
  Ed Huai hsin Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei.
  Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416, 2022.
[Clark et al., 2019] Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning. Electra:
  Pre-training text encoders as discriminators rather than generators. In Proceedings of International
  Conference on Learning Representations, 2019.
[Cobbe et al., 2021] Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz
  Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John
  Schulman. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
[Conneau et al., 2020] Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guil-
  laume Wenzek, Francisco Guzmán, Édouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov.
  Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting
  of the Association for Computational Linguistics, pages 8440–8451, 2020.
[Coste et al., 2024] Thomas Coste, Usman Anwar, Robert Kirk, and David Krueger. Reward model ensem-
  bles help mitigate overoptimization. In The Twelfth International Conference on Learning Representa-
  tions, 2024.
[Cui et al., 2024] Ganqu Cui, Lifan Yuan, Ning Ding, Guanming Yao, Bingxiang He, Wei Zhu, Yuan
  Ni, Guotong Xie, Ruobing Xie, Yankai Lin, Zhiyuan Liu, and Maosong Sun. ULTRAFEEDBACK:
  Boosting language models with scaled AI feedback. In Proceedings of the 41st International Conference
  on Machine Learning, volume 235, pages 9722–9744, 2024.
[Dai et al., 2023] Damai Dai, Yutao Sun, Li Dong, Yaru Hao, Shuming Ma, Zhifang Sui, and Furu Wei.
  Why can gpt learn in-context? language models secretly perform gradient descent as meta-optimizers.
  In Findings of the Association for Computational Linguistics: ACL 2023, pages 4005–4019, 2023.
[Dai et al., 2019] Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G Carbonell, Quoc Le, and Ruslan
  Salakhutdinov. Transformer-xl: Attentive language models beyond a fixed-length context. In Proceed-
  ings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2978–2988,
  2019.
[Dao et al., 2022] Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. Flashattention: Fast
  and memory-efficient exact attention with io-awareness. Advances in Neural Information Processing

  Systems, 35:16344–16359, 2022.
[Dehghani et al., 2018] Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz
  Kaiser. Universal transformers. arXiv preprint arXiv:1807.03819, 2018.
[Deletang et al., 2024] Gregoire Deletang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt, Tim
  Genewein, Christopher Mattern, Jordi Grau-Moya, Li Kevin Wenliang, Matthew Aitchison, Laurent
  Orseau, Marcus Hutter, and Joel Veness. Language modeling is compression. In The Twelfth Interna-
  tional Conference on Learning Representations, 2024.
[Deng et al., 2022] Mingkai Deng, Jianyu Wang, Cheng-Ping Hsieh, Yihan Wang, Han Guo, Tianmin Shu,
  Meng Song, Eric Xing, and Zhiting Hu. Rlprompt: Optimizing discrete text prompts with reinforcement
  learning. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,
  pages 3369–3391, 2022.
[Devlin et al., 2019] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-
  training of deep bidirectional transformers for language understanding. In Proceedings of the 2019
  Conference of the North American Chapter of the Association for Computational Linguistics: Human
  Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, 2019.
[Ding et al., 2024] Yiran Ding, Li Lyna Zhang, Chengruidong Zhang, Yuanyuan Xu, Ning Shang, Jiahang
  Xu, Fan Yang, and Mao Yang. Longrope: Extending llm context window beyond 2 million tokens. arXiv
  preprint arXiv:2402.13753, 2024.
[Dolan and Brockett, 2005] Bill Dolan and Chris Brockett. Automatically constructing a corpus of senten-
  tial paraphrases. In Proceedings of Third International Workshop on Paraphrasing (IWP2005), 2005.
[Dong et al., 2019] Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng
  Gao, Ming Zhou, and Hsiao-Wuen Hon. Unified language model pre-training for natural language
  understanding and generation. Advances in neural information processing systems, 32, 2019.
[Dong et al., 2022] Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun,
  Jingjing Xu, and Zhifang Sui. A survey on in-context learning. arXiv preprint arXiv:2301.00234, 2022.
[Dong et al., 2021] Yihe Dong, Jean-Baptiste Cordonnier, and Andreas Loukas. Attention is not all you
  need: Pure attention loses rank doubly exponentially with depth. In International Conference on Machine
  Learning, pages 2793–2803. PMLR, 2021.
[Drozdov et al., 2022] Andrew Drozdov, Nathanael Schärli, Ekin Akyürek, Nathan Scales, Xinying Song,
  Xinyun Chen, Olivier Bousquet, and Denny Zhou. Compositional semantic parsing with large language
  models. In Proceedings of The Eleventh International Conference on Learning Representations, 2022.
[Dua et al., 2022] Dheeru Dua, Shivanshu Gupta, Sameer Singh, and Matt Gardner. Successive prompting
  for decomposing complex questions. In Proceedings of the 2022 Conference on Empirical Methods in
  Natural Language Processing, pages 1251–1265, 2022.
[Dubey et al., 2024] Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-
  Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The llama 3 herd of
  models. arXiv preprint arXiv:2407.21783, 2024.
[Dubois et al., 2024] Yann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy
  Ba, Carlos Guestrin, Percy S Liang, and Tatsunori B Hashimoto. Alpacafarm: A simulation framework
  for methods that learn from human feedback. Advances in Neural Information Processing Systems, 36,
  2024.
[Eisenstein et al., 2023] Jacob Eisenstein, Chirag Nagpal, Alekh Agarwal, Ahmad Beirami, Alex D’Amour,
  DJ Dvijotham, Adam Fisch, Katherine Heller, Stephen Pfohl, Deepak Ramachandran, and Peter Shaw.
  Helping or herding? reward model ensembles mitigate but do not eliminate reward hacking. arXiv
  preprint arXiv:2312.09244, 2023.
[Elsken et al., 2019] Thomas Elsken, Jan Hendrik Metzen, and Frank Hutter. Neural architecture search:
  A survey. Journal of Machine Learning Research, 20(55):1–21, 2019.

[Erhan et al., 2010] Dumitru Erhan, Aaron Courville, Yoshua Bengio, and Pascal Vincent. Why does
  unsupervised pre-training help deep learning? In Proceedings of the thirteenth international conference
  on artificial intelligence and statistics, pages 201–208, 2010.
[Fan et al., 2019] Angela Fan, Edouard Grave, and Armand Joulin. Reducing transformer depth on demand
  with structured dropout. In Proceedings of International Conference on Learning Representations, 2019.
[Fedus et al., 2022] William Fedus, Barret Zoph, and Noam Shazeer. Switch transformers: Scaling to
  trillion parameter models with simple and efficient sparsity. The Journal of Machine Learning Research,
  23(1):5232–5270, 2022.
[Fernandes et al., 2023] Patrick Fernandes, Aman Madaan, Emmy Liu, António Farinhas, Pedro Henrique
  Martins, Amanda Bertsch, José G. C. de Souza, Shuyan Zhou, Tongshuang Wu, Graham Neubig, and
  André F. T. Martins. Bridging the gap: A survey on integrating (human) feedback for natural language
  generation. Transactions of the Association for Computational Linguistics, 11:1643–1668, 2023.
[Franklin and Graesser, 1996] Stan Franklin and Art Graesser. Is it an agent, or just a program?: A taxon-
   omy for autonomous agents. In International workshop on agent theories, architectures, and languages,
   pages 21–35. Springer, 1996.
[Frensch and Funke, 2014] Peter A Frensch and Joachim Funke. Complex problem solving: The European
   perspective. Psychology Press, 2014.
[Gale et al., 2019] Trevor Gale, Erich Elsen, and Sara Hooker. The state of sparsity in deep neural networks.
  arXiv preprint arXiv:1902.09574, 2019.
[Ganguli et al., 2023] Deep Ganguli, Amanda Askell, Nicholas Schiefer, Thomas I. Liao, Kamile Luko-
  siute, Anna Chen, Anna Goldie, Azalia Mirhoseini, Catherine Olsson, Danny Hernandez, Dawn Drain,
  Dustin Li, Eli Tran-Johnson, Ethan Perez, Jackson Kernion, Jamie Kerr, Jared Mueller, Joshua Landau,
  Kamal Ndousse, Karina Nguyen, Liane Lovitt, Michael Sellitto, Nelson Elhage, Noemí Mercado, Nova
  DasSarma, Oliver Rausch, Robert Lasenby, Robin Larson, Sam Ringer, Sandipan Kundu, Saurav Kada-
  vath, Scott Johnston, Shauna Kravec, Sheer El Showk, Tamera Lanham, Timothy Telleen-Lawton, Tom
  Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph,
  Sam McCandlish, Tom Brown, Christopher Olah, Jack Clark, Samuel R. Bowman, and Jared Kaplan.
  The capacity for moral self-correction in large language models. arXiv preprint arXiv:2302.07459, 2023.
[Gao et al., 2023] Leo Gao, John Schulman, and Jacob Hilton. Scaling laws for reward model overopti-
  mization. In International Conference on Machine Learning, pages 10835–10866. PMLR, 2023a.
[Gao et al., 2023] Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie
  Callan, and Graham Neubig. Pal: Program-aided language models. In International Conference on
  Machine Learning, pages 10764–10799. PMLR, 2023b.
[Gao et al., 2023] Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei
  Sun, and Haofen Wang. Retrieval-augmented generation for large language models: A survey. arXiv
  preprint arXiv:2312.10997, 2023c.
[Garg et al., 2022] Shivam Garg, Dimitris Tsipras, Percy S Liang, and Gregory Valiant. What can trans-
  formers learn in-context? a case study of simple function classes. Advances in Neural Information
  Processing Systems, 35:30583–30598, 2022.
[Ge et al., 2024] Yuan Ge, Yilun Liu, Chi Hu, Weibin Meng, Shimin Tao, Xiaofeng Zhao, Hongxia
  Ma, Li Zhang, Boxing Chen, Hao Yang, Bei Li, Tong Xiao, and Jingbo Zhu. Clustering and rank-
  ing: Diversity-preserved instruction selection through expert-aligned quality estimation. arXiv preprint
  arXiv:2402.18191, 2024.
[Gemma Team, 2024] Google DeepMind Gemma Team. Gemma: Open Models Based on Gemini Re-
  search and Technology, 2024.
[Goodhart, 1984] Charles AE Goodhart. Problems of monetary management: the UK experience. Springer,
  1984.

[Gordon et al., 2021] Mitchell A Gordon, Kevin Duh, and Jared Kaplan. Data and parameter scaling laws
  for neural machine translation. In Proceedings of the 2021 Conference on Empirical Methods in Natural
  Language Processing, pages 5915–5922, 2021.
[Gu and Dao, 2023] Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state
  spaces. arXiv preprint arXiv:2312.00752, 2023.
[Gunasekar et al., 2023] Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del
  Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, Adil
  Salim, Shital Shah, Harkirat Singh Behl, Xin Wang, Sébastien Bubeck, Ronen Eldan, Adam Tauman
  Kalai, Yin Tat Lee, and Yuanzhi Li. Textbooks are all you need. arXiv preprint arXiv:2306.11644, 2023.
[Guo et al., 2024] Qingyan Guo, Rui Wang, Junliang Guo, Bei Li, Kaitao Song, Xu Tan, Guoqing Liu, Jiang
  Bian, and Yujiu Yang. Connecting large language models with evolutionary algorithms yields powerful
  prompt optimizers. In The Twelfth International Conference on Learning Representations, 2024.
[Gupta and Berant, 2020] Ankit Gupta and Jonathan Berant. Gmat: Global memory augmentation for
  transformers. arXiv preprint arXiv:2006.03274, 2020.
[Gupta et al., 2021] Ankit Gupta, Guy Dar, Shaya Goodman, David Ciprut, and Jonathan Berant. Memory-
  efficient transformers via top-k attention. In Proceedings of the Second Workshop on Simple and Efficient
  Natural Language Processing, pages 39–52, 2021.
[Han et al., 2021] Xu Han, Zhengyan Zhang, Ning Ding, Yuxian Gu, Xiao Liu, Yuqi Huo, Jiezhong Qiu,
  Liang Zhang, Wentao Han, Minlie Huang, Qin Jin, Yanyan Lan, Yang Liu, Zhiyuan Liu, Zhiwu Lu,
  Xipeng Qiu, Ruihua Song, Jie Tang, Ji-Rong Wen, Jinhui Yuan, Wayne Xin Zhao, and Jun Zhu. Pre-
  trained models: Past, present and future. AI Open, 2:225–250, 2021.
[Han et al., 2024] Zeyu Han, Chao Gao, Jinyang Liu, Jeff Zhang, and Sai Qian Zhang. Parameter-efficient
  fine-tuning for large models: A comprehensive survey. arXiv preprint arXiv:2403.14608, 2024.
[Harlap et al., 2018] Aaron Harlap, Deepak Narayanan, Amar Phanishayee, Vivek Seshadri, Nikhil Deva-
  nur, Greg Ganger, and Phil Gibbons. Pipedream: Fast and efficient pipeline parallel dnn training. arXiv
  preprint arXiv:1806.03377, 2018.
[He et al., 2019] Kaiming He, Ross Girshick, and Piotr Dollár. Rethinking imagenet pre-training. In
  Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4918–4927, 2019.
[He et al., 2021] Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. Deberta: Decoding-
  enhanced bert with disentangled attention. In Proceedings of International Conference on Learning
  Representations, 2021.
[Hendrycks and Gimpel, 2016] Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus).
  arXiv preprint arXiv:1606.08415, 2016.
[Hendrycks et al., 2020] Dan Hendrycks, Xiaoyuan Liu, Eric Wallace, Adam Dziedzic, Rishabh Krishnan,
  and Dawn Song. Pretrained transformers improve out-of-distribution robustness. In Proceedings of the
  58th Annual Meeting of the Association for Computational Linguistics, pages 2744–2751, 2020.
[Hendrycks et al., 2021] Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn
  Song, and Jacob Steinhardt. Measuring massive multitask language understanding. In Proceedings of
  International Conference on Learning Representations, 2021.
[Hestness et al., 2017] Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun,
  Hassan Kianinejad, Md Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou. Deep learning scaling is
  predictable, empirically. arXiv preprint arXiv:1712.00409, 2017.
[Hewitt, 2024] John Hewitt. Instruction following without instruction tuning, 2024. URL https://nlp.
  stanford.edu/~johnhew/instruction-following.html.
[Hewitt et al., 2024] John Hewitt, Nelson F Liu, Percy Liang, and Christopher D Manning. Instruction
  following without instruction tuning. arXiv preprint arXiv:2409.14254, 2024.

[Hochreiter and Schmidhuber, 1997] Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory.
  Neural computation, 9(8):1735–1780, 1997.
[Hoffmann et al., 2022] Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor
  Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom
  Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Si-
  mon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and Laurent Sifre. Training
  compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022.
[Honovich et al., 2023] Or Honovich, Thomas Scialom, Omer Levy, and Timo Schick. Unnatural instruc-
  tions: Tuning language models with (almost) no human labor. In Proceedings of the 61st Annual Meeting
  of the Association for Computational Linguistics (Volume 1: Long Papers), pages 14409–14428, 2023.
[Houlsby et al., 2019] Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin
  De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. Parameter-efficient transfer
  learning for NLP. In Proceedings of the 36th International Conference on Machine Learning, pages
  2790–2799. PMLR, 2019.
[Hu et al., 2022] Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang,
  Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. In International
  Conference on Learning Representations, 2022.
[Huang, 2009] Liang Huang. Dynamic programming-based search algorithms in NLP. In Proceedings
  of Human Language Technologies: The 2009 Annual Conference of the North American Chapter of the
  Association for Computational Linguistics, Companion Volume: Tutorial Abstracts, 2009.
[Huang et al., 2019] Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Mia Xu Chen, Dehao
  Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, and Zhifeng Chen. Gpipe: Efficient
  training of giant neural networks using pipeline parallelism. Advances in neural information processing
  systems, 32, 2019.
[Hutchins et al., 2022] DeLesley Hutchins, Imanol Schlag, Yuhuai Wu, Ethan Dyer, and Behnam
  Neyshabur. Block-recurrent transformers. Advances in neural information processing systems, 35:
  33248–33261, 2022.
[Jelinek, 1998] Frederick Jelinek. Statistical methods for speech recognition. MIT Press, 1998.
[Jiang et al., 2023] Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Deven-
   dra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile
   Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril,
   Thomas Wang, Timothée Lacroix, and William El Sayed. Mistral 7b. arXiv preprint arXiv:2310.06825,
   2023a.
[Jiang et al., 2023] Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. Llmlingua:
   Compressing prompts for accelerated inference of large language models. In Proceedings of the 2023
   Conference on Empirical Methods in Natural Language Processing, pages 13358–13376, 2023b.
[Jiang et al., 2020] Zhengbao Jiang, Frank F Xu, Jun Araki, and Graham Neubig. How can we know what
   language models know? Transactions of the Association for Computational Linguistics, 8:423–438,
   2020.
[Jiao et al., 2020] Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang,
   and Qun Liu. Tinybert: Distilling bert for natural language understanding. In Findings of the Association
   for Computational Linguistics: EMNLP 2020, pages 4163–4174, 2020.
[Joshi et al., 2017] Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. Triviaqa: A large
   scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th
   Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1601–
   1611, 2017.
[Joshi et al., 2020] Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S Weld, Luke Zettlemoyer, and Omer
   Levy. Spanbert: Improving pre-training by representing and predicting spans. Transactions of the

  association for computational linguistics, 8:64–77, 2020.
[Jurafsky and Martin, 2008] Dan Jurafsky and James H. Martin. Speech and Language Processing (2nd
   ed.). Prentice Hall, 2008.
[Kahneman, 2011] Daniel Kahneman. Thinking, fast and slow. macmillan, 2011.
[Kaplan et al., 2020] Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Re-
  won Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language
  models. arXiv preprint arXiv:2001.08361, 2020.
[Katharopoulos et al., 2020] Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret.
  Transformers are rnns: Fast autoregressive transformers with linear attention. In International conference
  on machine learning, pages 5156–5165. PMLR, 2020.
[Khandelwal et al., 2020] Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike
  Lewis. Generalization through memorization: Nearest neighbor language models. In International
  Conference on Learning Representations, 2020.
[Khot et al., 2023] Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark,
  and Ashish Sabharwal. Decomposed prompting: A modular approach for solving complex tasks. In
  Proceedings of The Eleventh International Conference on Learning Representations, 2023.
[Kim et al., 2023] Sehoon Kim, Coleman Hooper, Thanakul Wattanawong, Minwoo Kang, Ruohan
  Yan, Hasan Genc, Grace Dinh, Qijing Huang, Kurt Keutzer, Michael W. Mahoney, Yakun Sophia
  Shao, and Amir Gholami. Full stack optimization of transformer inference: a survey. arXiv preprint
  arXiv:2302.14017, 2023.
[Kirkpatrick et al., 2017] James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume
  Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska,
  Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell. Overcoming catastrophic
  forgetting in neural networks. Proceedings of the national academy of sciences, 114(13):3521–3526,
  2017.
[Koehn, 2010] Philipp Koehn. Statistical Machine Translation. Cambridge University Press, 2010.
[Kojima et al., 2022] Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke
  Iwasawa. Large language models are zero-shot reasoners. Advances in neural information processing
  systems, 35:22199–22213, 2022.
[Korthikanti et al., 2023] Vijay Anand Korthikanti, Jared Casper, Sangkug Lym, Lawrence McAfee,
  Michael Andersch, Mohammad Shoeybi, and Bryan Catanzaro. Reducing activation recomputation in
  large transformer models. Proceedings of Machine Learning and Systems, 5, 2023.
[Krakovna et al., 2020] Victoria Krakovna, Jonathan Uesato, Vladimir Mikulik, Matthew
  Rahtz, Tom Everitt, Ramana Kumar, Zac Kenton, Jan Leike, and Shane Legg.      Specifi-
  cation gaming: the flip side of ai ingenuity. https://deepmind.google/discover/blog/
  specification-gaming-the-flip-side-of-ai-ingenuity, 2020.
[Kung and Peng, 2023] Po-Nien Kung and Nanyun Peng. Do models really learn to follow instructions?
  an empirical study of instruction tuning. arXiv preprint arXiv:2305.11383, 2023.
[Kwon et al., 2023] Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao
  Yu, Joseph E Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language
  model serving with pagedattention. arXiv preprint arXiv:2309.06180, 2023.
[Lake and Baroni, 2018] Brenden Lake and Marco Baroni. Generalization without systematicity: On
  the compositional skills of sequence-to-sequence recurrent networks. In International conference on
  machine learning, pages 2873–2882. PMLR, 2018.
[Lambert et al., 2024] Nathan Lambert, Valentina Pyatkin, Jacob Morrison, LJ Miranda, Bill Yuchen
  Lin, Khyathi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, Noah A. Smith, and Han-
  naneh Hajishirzi. Rewardbench: Evaluating reward models for language modeling. arXiv preprint

  arXiv:2403.13787, 2024.
[Lample and Conneau, 2019] Guillaume Lample and Alexis Conneau. Cross-lingual language model
  pretraining. arXiv preprint arXiv:1901.07291, 2019.
[Lan et al., 2020] Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and
  Radu Soricut. Albert: A lite bert for self-supervised learning of language representations. In Proceedings
  of International Conference on Learning Representations, 2020.
[Lee et al., 2023] Harrison Lee, Samrat Phatale, Hassan Mansoor, Kellie Ren Lu, Thomas Mesnard, Johan
  Ferret, Colton Bishop, Ethan Hall, Victor Carbune, and Abhinav Rastogi. Rlaif: Scaling reinforcement
  learning from human feedback with ai feedback. arXiv preprint arXiv:2309.00267, 2023.
[Lester et al., 2021] Brian Lester, Rami Al-Rfou, and Noah Constant. The power of scale for parameter-
  efficient prompt tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Lan-
  guage Processing, pages 3045–3059, 2021.
[Lewis et al., 2020] Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mo-
  hamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. Bart: Denoising sequence-to-sequence
  pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th
  Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, 2020.
[Li et al., 2023] Bei Li, Rui Wang, Junliang Guo, Kaitao Song, Xu Tan, Hany Hassan, Arul Menezes, Tong
  Xiao, Jiang Bian, and JingBo Zhu. Deliberate then generate: Enhanced prompting framework for text
  generation. arXiv preprint arXiv:2305.19835, 2023a.
[Li, 2011] Hang Li. Learning to Rank for Information Retrieval and Natural Language Processing. Online
  access: Morgan & Claypool Synthesis Collection Five. Morgan & Claypool Publishers, 2011. ISBN
  9781608457076.
[Li et al., 2022] Huayang Li, Yixuan Su, Deng Cai, Yan Wang, and Lemao Liu. A survey on retrieval-
  augmented text generation. arXiv preprint arXiv:2202.01110, 2022.
[Li et al., 2024] Shanda Li, Chong You, Guru Guruganesh, Joshua Ainslie, Santiago Ontanon, Manzil
  Zaheer, Sumit Sanghai, Yiming Yang, Sanjiv Kumar, and Srinadh Bhojanapalli. Functional interpolation
  for relative positions improves long context transformers. In The Twelfth International Conference on
  Learning Representations, 2024.
[Li et al., 2023] Shenggui Li, Fuzhao Xue, Chaitanya Baranwal, Yongbin Li, and Yang You. Sequence
  parallelism: Long sequence training from system perspective. In Proceedings of the 61st Annual Meeting
  of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2391–2404, 2023b.
[Li and Liang, 2021] Xiang Lisa Li and Percy Liang. Prefix-tuning: Optimizing continuous prompts for
  generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics
  and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers),
  pages 4582–4597, 2021.
[Li, 2023] Yinheng Li. A practical survey on zero-shot prompt design for in-context learning. In Proceed-
  ings of the 14th International Conference on Recent Advances in Natural Language Processing, pages
  641–647, 2023.
[Li et al., 2023] Yucheng Li, Bo Dong, Frank Guerin, and Chenghua Lin. Compressing context to enhance
  inference efficiency of large language models. In Proceedings of the 2023 Conference on Empirical
  Methods in Natural Language Processing, pages 6342–6353, 2023c.
[Lialin et al., 2023] Vladislav Lialin, Vijeta Deshpande, and Anna Rumshisky. Scaling down to scale up:
  A guide to parameter-efficient fine-tuning. arXiv preprint arXiv:2303.15647, 2023.
[Lightman et al., 2024] Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker,
  Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. Let’s verify step by step. In The
  Twelfth International Conference on Learning Representations, 2024.
[Liu et al., 2024] Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang

  Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. Deepseek-v3 technical report. arXiv preprint
  arXiv:2412.19437, 2024a.
[Liu et al., 2022] Jiachang Liu, Dinghan Shen, Yizhe Zhang, William B Dolan, Lawrence Carin, and
  Weizhu Chen. What makes good in-context examples for gpt-3? In Proceedings of Deep Learning Inside
  Out (DeeLIO 2022): The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning
  Architectures, pages 100–114, 2022.
[Liu et al., 2023] Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham
  Neubig. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language
  processing. ACM Computing Surveys, 55(9):1–35, 2023a.
[Liu et al., 2024] Tianqi Liu, Yao Zhao, Rishabh Joshi, Misha Khalman, Mohammad Saleh, Peter J Liu, and
  Jialu Liu. Statistical rejection sampling improves preference optimization. In The Twelfth International
  Conference on Learning Representations, 2024b.
[Liu, 2009] Tie-Yan Liu. Learning to rank for information retrieval. Foundations and Trends® in Informa-
  tion Retrieval, 3(3):225–331, 2009.
[Liu et al., 2023] Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie
  Tang. Gpt understands, too. AI Open, 2023b.
[Liu et al., 2023] Xiaoxia Liu, Jingyi Wang, Jun Sun, Xiaohan Yuan, Guoliang Dong, Peng Di, Wenhai
  Wang, and Dongxia Wang. Prompting frameworks for large language models: A survey. arXiv preprint
  arXiv:2311.12785, 2023c.
[Liu et al., 2024] Xinyu Liu, Runsong Zhao, Pengcheng Huang, Chunyang Xiao, Bei Li, Jingang Wang,
  Tong Xiao, and Jingbo Zhu. Forgetting curve: A reliable method for evaluating memorization capability
  for long-context models. In Proceedings of the 2024 Conference on Empirical Methods in Natural
  Language Processing, pages 4667–4682, 2024c.
[Liu et al., 2019] Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy,
  Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining
  approach. arXiv preprint arXiv:1907.11692, 2019.
[Longpre et al., 2023] Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay,
  Denny Zhou, Quoc V. Le, Barret Zoph, Jason Wei, and Adam Roberts. The flan collection: Designing
  data and methods for effective instruction tuning. In International Conference on Machine Learning,
  pages 22631–22648. PMLR, 2023.
[Ma et al., 2023] Xuezhe Ma, Chunting Zhou, Xiang Kong, Junxian He, Liangke Gui, Graham Neubig,
  Jonathan May, and Luke Zettlemoyer. Mega: Moving average equipped gated attention. In The Eleventh
  International Conference on Learning Representations, 2023.
[Ma et al., 2024] Xuezhe Ma, Xiaomeng Yang, Wenhan Xiong, Beidi Chen, Lili Yu, Hao Zhang, Jonathan
  May, Luke Zettlemoyer, Omer Levy, and Chunting Zhou. Megalodon: Efficient llm pretraining and
  inference with unlimited context length. arXiv preprint arXiv:2404.08801, 2024.
[Madaan et al., 2024] Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao,
  Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bod-
  hisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark.
  Self-refine: Iterative refinement with self-feedback. Advances in Neural Information Processing Sys-
  tems, 36, 2024.
[Manning, 2022] Christopher D Manning. Human language understanding & reasoning. Daedalus, 151
  (2):127–138, 2022.
[Marcus, 1993] Gary F Marcus. Negative evidence in language acquisition. Cognition, 46(1):53–85, 1993.
[Martins et al., 2022] Pedro Henrique Martins, Zita Marinho, and André FT Martins. ∞-former: Infinite
  memory transformer-former: Infinite memory transformer. In Proceedings of the 60th Annual Meeting
  of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5468–5485, 2022.

امتیاز کاربران به این مقاله

☆☆☆☆☆

0 نفر امتیاز داده اند. میانگین: 0.0 از 5

 

0 نظر

نظر محترم شما در مورد مقاله های وب سایت برنامه نویسی و پایگاه داده

نظرات محترم شما در خدمات رسانی بهتر ما را یاری می نمایند. لطفا اگر مایل بودید یک نظر ما را مهمان فرمائید. آدرس ایمیل و وب سایت شما نمایش داده نخواهد شد.

0 / 500

اطلاعات تماس

  • آدرس:اصفهان-خیابان ام کلثوم غربی - بعد خیابان تخم چی - بیست متر بعد از پیتزا ننه شب - کوچه تعمیر گاه سمار زغالی - پلاک 354 - درب مشکی - طبقه هفتم
  • آدرس ایمیل:najafzade@gmail.com
  • وب سایت:http://www.a00b.com/
  • تلفن ثابت:(+98)9131253620
  • تلفن همراه:09131253620