منابع و نمایهٔ کتاب Foundations of Large Language Models — بخش دوم

منابع و نمایهٔ کتاب Foundations of Large Language Models — بخش دوم

توسط admin | گروه هوش مصنوعی | 1405/05/18

نظرات 0

منابع و نمایهٔ کتاب Foundations of Large Language Models — بخش دوم

عنوان اصلی کتاب: Foundations of Large Language Models

نویسندگان: Tong Xiao و Jingbo Zhu

سازمان: NLP Lab, Northeastern University & NiuTrans Research

زبان اصلی: انگلیسی

بازهٔ منبع: صفحات PDF 221 تا 231

مجوز منبع: Creative Commons Attribution-NonCommercial 4.0 (CC BY-NC 4.0)

تاریخ ترجمه: 1405/05/18 / 2026-08-09

اعتبار ترجمه: ترجمه با کمک هوش مصنوعی

منابع کتاب — بخش دوم

ادامهٔ فهرست منابع، با اطلاعات کتاب‌شناختی اصلی و بدون تغییر در شناسه‌ها و نام آثار.

[Mavi et al., 2024] Vaibhav Mavi, Anubhav Jangra, and Adam Jatowt. Multi-hop question answering.
  Foundations and Trends® in Information Retrieval, 17(5):457–586, 2024.
[Michel et al., 2019] Paul Michel, Omer Levy, and Graham Neubig. Are sixteen heads really better than
  one? Advances in neural information processing systems, 32, 2019.
[Micikevicius et al., 2018] Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich
  Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao
  Wu. Mixed precision training. In Proceedings of International Conference on Learning Representations,
  2018.
[Miettinen, 1999] Kaisa Miettinen. Nonlinear multiobjective optimization, volume 12. Springer Science
  & Business Media, 1999.
[Mikolov et al., 2013] Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. Efficient estimation
  of word representations in vector space. In Proceedings of the International Conference on Learning
  Representations (ICLR 2013), 2013a.
[Mikolov et al., 2013] Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, and Jeffrey Dean. Dis-
  tributed representations of words and phrases and their compositionality. In Proceedings of the 26th In-
  ternational Conference on Neural Information Processing Systems - Volume 2, pages 3111–3119, 2013b.
[Min et al., 2019] Sewon Min, Victor Zhong, Luke Zettlemoyer, and Hannaneh Hajishirzi. Multi-hop read-
  ing comprehension through question decomposition and rescoring. In Proceedings of the 57th Annual
  Meeting of the Association for Computational Linguistics, pages 6097–6109, 2019.
[Minaee et al., 2024] Shervin Minaee, Tomas Mikolov, Narjes Nikzad, Meysam Chenaghlu, Richard
  Socher, Xavier Amatriain, and Jianfeng Gao. Large language models: A survey. arXiv preprint
  arXiv:2402.06196, 2024.
[Mishra et al., 2022] Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. Cross-
  task generalization via natural language crowdsourcing instructions. In Proceedings of the 60th Annual
  Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3470–3487,
  2022.
[Mnih et al., 2016] Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Tim Harley,
  Timothy P Lillicrap, David Silver, and Koray Kavukcuoglu. Asynchronous methods for deep reinforce-
  ment learning. In Proceedings of the 33rd International Conference on International Conference on
  Machine Learning, pages 1928–1937, 2016.
[Mohtashami and Jaggi, 2024] Amirkeivan Mohtashami and Martin Jaggi. Random-access infinite context
  length for transformers. Advances in Neural Information Processing Systems, 36, 2024.
[Mu et al., 2024] Jesse Mu, Xiang Li, and Noah Goodman. Learning to compress prompts with gist tokens.
  Advances in Neural Information Processing Systems, 36, 2024.
[Munkhdalai et al., 2024] Tsendsuren Munkhdalai, Manaal Faruqui, and Siddharth Gopal. Leave no context
  behind: Efficient infinite context transformers with infini-attention. arXiv preprint arXiv:2404.07143,
  2024.
[Nakano et al., 2021] Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina
  Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna
  Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman.
  Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332,
  2021.
[Narayanan et al., 2021] Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley,
  Mostofa Patwary, Vijay Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan
  Catanzaro, Amar Phanishayee, and Matei Zaharia. Efficient large-scale language model training on
  gpu clusters using megatron-lm. In Proceedings of the International Conference for High Performance
  Computing, Networking, Storage and Analysis, pages 1–15, 2021.

[Ng et al., 1999] Andrew Y Ng, Daishi Harada, and Stuart J Russell. Policy invariance under reward
  transformations: Theory and application to reward shaping. In Proceedings of the Sixteenth International
  Conference on Machine Learning, pages 278–287, 1999.
[OpenAI, 2024] OpenAI. Learning to reason with llms, September 2024. URL https://openai.com/
  index/learning-to-reason-with-llms/.
[Ouyang et al., 2022] Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela
  Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton,
  Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan
  Leike, and Ryan Lowe. Training language models to follow instructions with human feedback. Advances
  in Neural Information Processing Systems, 35:27730–27744, 2022.
[Pal et al., 2023] Koyena Pal, Jiuding Sun, Andrew Yuan, Byron C Wallace, and David Bau. Future lens:
  Anticipating subsequent tokens from a single hidden state. In Proceedings of the 27th Conference on
  Computational Natural Language Learning (CoNLL), pages 548–560, 2023.
[Pan et al., 2022] Alexander Pan, Kush Bhatia, and Jacob Steinhardt. The effects of reward misspecifica-
  tion: Mapping and mitigating misaligned models. In International Conference on Learning Representa-
  tions, 2022.
[Pan et al., 2024] Liangming Pan, Michael Saxon, Wenda Xu, Deepak Nathani, Xinyi Wang, and
  William Yang Wang. Automatically correcting large language models: Surveying the landscape of
  diverse automated correction strategies. Transactions of the Association for Computational Linguistics,
  12:484–506, 2024.
[Parisi et al., 2022] Aaron Parisi, Yao Zhao, and Noah Fiedel. Talm: Tool augmented language models.
  arXiv preprint arXiv:2205.12255, 2022.
[Parisi et al., 2019] German I Parisi, Ronald Kemker, Jose L Part, Christopher Kanan, and Stefan Wermter.
  Continual lifelong learning with neural networks: A review. Neural networks, 113:54–71, 2019.
[Parmar et al., 2018] Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer,
  Alexander Ku, and Dustin Tran. Image transformer. In International conference on machine learn-
  ing, pages 4055–4064. PMLR, 2018.
[Penedo et al., 2023] Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessan-
  dro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. The refined-
  web dataset for falcon llm: outperforming curated corpora with web data, and web data only. arXiv
  preprint arXiv:2306.01116, 2023.
[Peng et al., 2024] Bowen Peng, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole. YaRN: Efficient con-
  text window extension of large language models. In The Twelfth International Conference on Learning
  Representations, 2024.
[Pennington et al., 2014] Jeffrey Pennington, Richard Socher, and Christopher D. Manning. Glove: Global
  vectors for word representation. In Proceedings of Empirical Methods in Natural Language Processing
  (EMNLP), pages 1532–1543, 2014.
[Peters et al., 2018] Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark,
  Kenton Lee, and Luke Zettlemoyer. Deep contextualized word representations. In Proceedings of the
  2018 Conference of the North American Chapter of the Association for Computational Linguistics: Hu-
  man Language Technologies, Volume 1 (Long Papers), 2018.
[Plackett, 1975] Robin L Plackett. The analysis of permutations. Journal of the Royal Statistical Society
   Series C: Applied Statistics, 24(2):193–202, 1975.
[Prasad et al., 2023] Archiki Prasad, Peter Hase, Xiang Zhou, and Mohit Bansal. Grips: Gradient-free, edit-
   based instruction search for prompting large language models. In Proceedings of the 17th Conference of
   the European Chapter of the Association for Computational Linguistics, pages 3845–3864, 2023.

[Press et al., 2022] Ofir Press, Noah Smith, and Mike Lewis. Train short, test long: Attention with lin-
   ear biases enables input length extrapolation. In Proceedings of International Conference on Learning
   Representations, 2022.
[Press et al., 2023] Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah A Smith, and Mike Lewis.
   Measuring and narrowing the compositionality gap in language models. In Findings of the Association
   for Computational Linguistics: EMNLP 2023, pages 5687–5711, 2023.
[Pryzant et al., 2023] Reid Pryzant, Dan Iter, Jerry Li, Yin Tat Lee, Chenguang Zhu, and Michael Zeng.
   Automatic prompt optimization with "gradient descent" and beam search. In The 2023 Conference on
   Empirical Methods in Natural Language Processing, 2023.
[Qiu et al., 2020] Xipeng Qiu, Tianxiang Sun, Yige Xu, Yunfan Shao, Ning Dai, and Xuanjing Huang.
  Pre-trained models for natural language processing: A survey. Science China Technological Sciences,
  63(10):1872–1897, 2020.
[Radford et al., 2018] Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving
  language understanding by generative pre-training. OpenAI, 2018.
[Radford et al., 2019] Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya
  Sutskever. Language models are unsupervised multitask learners. OpenAI blog, 1(8), 2019.
[Radford et al., 2021] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sand-
  hini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya
  Sutskever. Learning transferable visual models from natural language supervision. In International
  conference on machine learning, pages 8748–8763. PMLR, 2021.
[Rae et al., 2019] Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, Chloe Hillier, and Timothy P
  Lillicrap. Compressive transformers for long-range sequence modelling. In International Conference on
  Learning Representations, 2019.
[Rafailov et al., 2024] Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano
  Ermon, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward
  model. Advances in Neural Information Processing Systems, 36, 2024.
[Raffel et al., 2020] Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael
  Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified
  text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020.
[Ramachandran et al., 2017] Prajit Ramachandran, Barret Zoph, and Quoc V Le. Searching for activation
  functions. arXiv preprint arXiv:1710.05941, 2017.
[Rolnick et al., 2019] David Rolnick, Arun Ahuja, Jonathan Schwarz, Timothy Lillicrap, and Gregory
  Wayne. Experience replay for continual learning. Advances in Neural Information Processing Systems,
  32, 2019.
[Rosenfeld et al., 2020] Jonathan S Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, and Nir Shavit. A con-
  structive prediction of the generalization error across scales. In Proceedings of International Conference
  on Learning Representations, 2020.
[Ruan et al., 2024] Junhao Ruan, Long Meng, Weiqiao Shan, Tong Xiao, and Jingbo Zhu. A survey of llm
  surveys. https://github.com/NiuTrans/ABigSurveyOfLLMs, 2024.
[Rubin et al., 2022] Ohad Rubin, Jonathan Herzig, and Jonathan Berant. Learning to retrieve prompts
  for in-context learning. In Proceedings of the 2022 Conference of the North American Chapter of the
  Association for Computational Linguistics: Human Language Technologies, pages 2655–2671, 2022.
[Russell, 2019] Stuart Russell. Human Compatible: Artificial Intelligence and the Problem of Controls.
  Viking, 2019.
[Sanh et al., 2020] Victor Sanh, Thomas Wolf, and Alexander Rush. Movement pruning: Adaptive sparsity
  by fine-tuning. Advances in Neural Information Processing Systems, 33:20378–20389, 2020.

[Sanh et al., 2022] Victor Sanh, Albert Webson, Colin Raffel, Stephen Bach, Lintang Sutawika, Zaid
  Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish
  Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak,
  Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen,
  Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht
  Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Bider-
  man, Leo Gao, Thomas Wolf, and Alexander M Rush. Multitask prompted training enables zero-shot
  task generalization. In Proceedings of International Conference on Learning Representations, 2022.
[Schick et al., 2023] Timo Schick, Jane A. Yu, Zhengbao Jiang, Fabio Petroni, Patrick Lewis, Gautier Izac-
  ard, Qingfei You, Christoforos Nalmpantis, Edouard Grave, and Sebastian Riedel. PEER: A collaborative
  language model. In Proceedings of The Eleventh International Conference on Learning Representations,
  2023.
[Schick et al., 2024] Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric
  Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. Toolformer: Language models can
  teach themselves to use tools. Advances in Neural Information Processing Systems, 36, 2024.
[Schmidhuber, 2015] Jürgen Schmidhuber. Deep learning in neural networks: An overview. Neural net-
  works, 61:85–117, 2015.
[Schulman et al., 2015] John Schulman, Sergey Levine, Philipp Moritz, Michael Jordan, and Pieter Abbeel.
  Trust region policy optimization. In Proceedings of the 32nd International Conference on International
  Conference on Machine Learning-Volume 37, pages 1889–1897, 2015.
[Schulman et al., 2017] John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov.
  Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
[Sennrich et al., 2016] Rico Sennrich, Barry Haddow, and Alexandra Birch. Improving neural machine
  translation models with monolingual data. In Proceedings of the 54th Annual Meeting of the Association
  for Computational Linguistics (Volume 1: Long Papers), pages 86–96, 2016.
[Seo et al., 2017] Minjoon Seo, Aniruddha Kembhavi, Ali Farhadi, and Hannaneh Hajishirzi. Bidirectional
  attention flow for machine comprehension. In Proceedings of International Conference on Learning
  Representations, 2017.
[Shannon, 1951] Claude E Shannon. Prediction and entropy of printed english. Bell system technical
  journal, 30(1):50–64, 1951.
[Shaw et al., 2018] Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. Self-attention with relative position
  representations. In Proceedings of the 2018 Conference of the North American Chapter of the Associ-
  ation for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages
  464–468, 2018.
[Shazeer, 2019] Noam Shazeer. Fast transformer decoding: One write-head is all you need. arXiv preprint
  arXiv:1911.02150, 2019.
[Shazeer, 2020] Noam Shazeer. Glu variants improve transformer. arXiv preprint arXiv:2002.05202, 2020.
[Shen et al., 2020] Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W
  Mahoney, and Kurt Keutzer. Q-bert: Hessian based ultra low precision quantization of bert. In Proceed-
  ings of the AAAI Conference on Artificial Intelligence, volume 34, pages 8815–8821, 2020.
[Shoeybi et al., 2019] Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper,
  and Bryan Catanzaro. Megatron-lm: Training multi-billion parameter language models using model
  parallelism. arXiv preprint arXiv:1909.08053, 2019.
[Skalse et al., 2022] Joar Skalse, Nikolaus Howe, Dmitrii Krasheninnikov, and David Krueger. Defining
  and characterizing reward gaming. Advances in Neural Information Processing Systems, 35:9460–9471,
  2022.
[Snell et al., 2022] Charlie Snell, Dan Klein, and Ruiqi Zhong. Learning by distilling context. arXiv

  preprint arXiv:2209.15189, 2022.
[Socher et al., 2013] Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning,
  Andrew Y Ng, and Christopher Potts. Recursive deep models for semantic compositionality over a
  sentiment treebank. In Proceedings of the 2013 conference on empirical methods in natural language
  processing, pages 1631–1642, 2013.
[Song et al., 2019] Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. Mass: Masked sequence
  to sequence pre-training for language generation. In International Conference on Machine Learning,
  pages 5926–5936. PMLR, 2019.
[Stiennon et al., 2020] Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea
   Voss, Alec Radford, Dario Amodei, and Paul F Christiano. Learning to summarize with human feedback.
   Advances in Neural Information Processing Systems, 33:3008–3021, 2020.
[Su et al., 2024] Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. Ro-
  former: Enhanced transformer with rotary position embedding. Neurocomputing, 568:127063, 2024.
[Sun et al., 2020] Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou.
  Mobilebert: a compact task-agnostic bert for resource-limited devices. In Proceedings of the 58th Annual
  Meeting of the Association for Computational Linguistics, pages 2158–2170, 2020.
[Sutskever et al., 2014] Ilya Sutskever, Oriol Vinyals, and Quoc V Le. Sequence to sequence learning with
  neural networks. Advances in neural information processing systems, 27, 2014.
[Sutton and Barto, 2018] Richard S. Sutton and Andrew G. Barto. Reinforcement Learning: An Introduc-
  tion (2nd ed.). The MIT Press, 2018.
[Szepesvári, 2010] Csaba Szepesvári. Algorithms for reinforcement learning. Synthesis Lectures on Arti-
  ficial Intelligence and Machine Learning, 4(1):1–103, 2010.
[Talmor and Berant, 2018] Alon Talmor and Jonathan Berant. The web as a knowledge-base for answering
  complex questions. arXiv preprint arXiv:1803.06643, 2018.
[Taori et al., 2023] Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos
  Guestrin, Percy Liang, and Tatsunori B. Hashimoto. Stanford alpaca: An instruction-following llama
  model. https://github.com/tatsu-lab/stanford_alpaca, 2023.
[Tay et al., 2020] Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. Efficient transformers: A
  survey. CoRR, abs/2009.06732, 2020.
[Team et al., 2024] Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin,
  Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, et al.
  Gemma 2: Improving open language models at a practical size. arXiv preprint arXiv:2408.00118, 2024.
[Teknium, 2023] Teknium. Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants,
  2023. URL https://huggingface.co/datasets/teknium/OpenHermes-2.5.
[Touvron et al., 2023] Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne
  Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Ro-
  driguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation
  language models. arXiv preprint arXiv:2302.13971, 2023a.
[Touvron et al., 2023] Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi,
  Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel,
  Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernan-
  des, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony
  Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Is-
  abel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee,
  Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor
  Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schel-
  ten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor,

  Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan,
  Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas
  Scialom. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288,
  2023b.
[Uesato et al., 2022] Jonathan Uesato, Nate Kushman, Ramana Kumar, Francis Song, Noah Siegel, Lisa
  Wang, Antonia Creswell, Geoffrey Irving, and Irina Higgins. Solving math word problems with process-
  and outcome-based feedback. arXiv preprint arXiv:2211.14275, 2022.
[Vaswani et al., 2017] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N
  Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Proceedings of Advances in
  Neural Information Processing Systems, volume 30, 2017.
[Von Oswald et al., 2023] Johannes Von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento,
  Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov. Transformers learn in-context by
  gradient descent. In Proceedings of International Conference on Machine Learning, pages 35151–
  35174. PMLR, 2023.
[Wang et al., 2024] Chenglong Wang, Hang Zhou, Yimin Hu, Yifu Huo, Bei Li, Tongran Liu, Tong Xiao,
  and Jingbo Zhu. Esrl: Efficient sampling-based reinforcement learning for sequence generation. In
  Proceedings of the AAAI Conference on Artificial Intelligence, pages 19107–19115, 2024.
[Wang et al., 2023] Liyuan Wang, Xingxing Zhang, Hang Su, and Jun Zhu. A comprehensive survey of
  continual learning: Theory, method and application. arXiv preprint arXiv:2302.00487, 2023a.
[Wang et al., 2019] Qiang Wang, Bei Li, Tong Xiao, Jingbo Zhu, Changliang Li, Derek F Wong, and
  Lidia S Chao. Learning deep transformer models for machine translation. In Proceedings of the 57th
  Annual Meeting of the Association for Computational Linguistics, pages 1810–1822, 2019.
[Wang et al., 2022] Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou.
  Rationale-augmented ensembles in language models. arXiv preprint arXiv:2207.00747, 2022a.
[Wang et al., 2023] Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H Chi, Sharan Narang,
  Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in lan-
  guage models. In Proceedings of The Eleventh International Conference on Learning Representations,
  2023b.
[Wang et al., 2022] Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza
  Mirzaei, Atharva Naik, Arjun Ashok, Arut Selvan Dhanasekaran, Anjana Arunkumar, David Stap, Es-
  haan Pathak, Giannis Karamanolakis, Haizhi Gary Lai, Ishan Purohit, Ishani Mondal, Jacob Anderson,
  Kirby Kuznia, Krima Doshi, Kuntal Kumar Pal, Maitreya Patel, Mehrad Moradshahi, Mihir Parmar, Mi-
  rali Purohit, Neeraj Varshney, Phani Rohitha Kaza, Pulkit Verma, Ravsehaj Singh Puri, Rushang Karia,
  Savan Doshi, Shailaja Keyur Sampat, Siddhartha Mishra, Sujan Reddy A, Sumanta Patro, Tanay Dixit,
  and Xudong Shen. Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp
  tasks. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,
  pages 5085–5109, 2022b.
[Wang et al., 2023] Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khy-
  athi Raghavi Chandu, David Wadden, Kelsey MacMillan, Noah A. Smith, Iz Beltagy, and Hannaneh
  Hajishirzi. How far can camels go? exploring the state of instruction tuning on open resources. Ad-
  vances in Neural Information Processing Systems, 36:74764–74786, 2023c.
[Wang et al., 2023] Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel
  Khashabi, and Hannaneh Hajishirzi. Self-instruct: Aligning language models with self-generated in-
  structions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics
  (Volume 1: Long Papers), pages 13484–13508, 2023d.
[Wang et al., 2023] Zhenyi Wang, Enneng Yang, Li Shen, and Heng Huang. A comprehensive survey of
  forgetting in deep learning beyond continual learning. arXiv preprint arXiv:2307.09218, 2023e.

[Warstadt et al., 2019] Alex Warstadt, Amanpreet Singh, and Samuel R Bowman. Neural network accept-
  ability judgments. Transactions of the Association for Computational Linguistics, 7:625–641, 2019.
[Wei et al., 2022] Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan
  Du, Andrew M Dai, and Quoc V Le. Finetuned language models are zero-shot learners. In Proceedings
  of International Conference on Learning Representations, 2022a.
[Wei et al., 2022] Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud,
  Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol
  Vinyals, Percy Liang, Jeff Dean, and William Fedus. Emergent abilities of large language models. arXiv
  preprint arXiv:2206.07682, 2022b.
[Wei et al., 2022] Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia,
  Ed H. Chi, Quoc V. Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language
  models. Advances in Neural Information Processing Systems, 35:24824–24837, 2022c.
[Welleck et al., 2023] Sean Welleck, Ximing Lu, Peter West, Faeze Brahman, Tianxiao Shen, Daniel
  Khashabi, and Yejin Choi. Generating sequences by learning to self-correct. In Proceedings of The
  Eleventh International Conference on Learning Representations, 2023.
[Weng, 2021] Lilian Weng. How to train really large models on many gpus? lilianweng.github.io, Sep
  2021. URL https://lilianweng.github.io/posts/2021-09-25-train-large/.
[Wiener, 1960] Norbert Wiener. Some moral and technical consequences of automation: As machines
  learn they may develop unforeseen strategies at rates that baffle their programmers. Science, 131(3410):
  1355–1358, 1960.
[Williams et al., 2018] Adina Williams, Nikita Nangia, and Samuel Bowman. A broad-coverage challenge
  corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North
  American Chapter of the Association for Computational Linguistics: Human Language Technologies,
  Volume 1 (Long Papers), pages 1112–1122, 2018.
[Williams, 1992] Ronald J Williams. Simple statistical gradient-following algorithms for connectionist
  reinforcement learning. Machine learning, 8:229–256, 1992.
[Wingate et al., 2022] David Wingate, Mohammad Shoeybi, and Taylor Sorensen. Prompt compression
  and contrastive conditioning for controllability and toxicity reduction in language models. In Findings
  of the Association for Computational Linguistics: EMNLP 2022, pages 5621–5634, 2022.
[Wu et al., 2024] Wilson Wu, John X Morris, and Lionel Levine. Do language models plan for future
  tokens? arXiv preprint arXiv:2404.00859, 2024.
[Wu et al., 2021] Yuhuai Wu, Markus Norman Rabe, DeLesley Hutchins, and Christian Szegedy. Memo-
  rizing transformers. In Proceedings of International Conference on Learning Representations, 2021.
[Wu et al., 2023] Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu,
  Noah A. Smith, Mari Ostendorf, and Hannaneh Hajishirzi. Fine-grained human feedback gives better
  rewards for language model training. In Thirty-seventh Conference on Neural Information Processing
  Systems, 2023.
[Xia et al., 2024] Mengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora, and Danqi Chen.
  Less: Selecting influential data for targeted instruction tuning. arXiv preprint arXiv:2402.04333, 2024.
[Xiao et al., 2024] Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. Efficient
  streaming language models with attention sinks. In Proceedings of The Twelfth International Conference
  on Learning Representations, 2024.
[Xiao and Zhu, 2023] Tong Xiao and Jingbo Zhu. Introduction to transformers: an nlp perspective. arXiv
  preprint arXiv:2311.17633, 2023.
[Xiao et al., 2013] Tong Xiao, Jingbo Zhu, and Tongran Liu. Bagging and boosting statistical machine
  translation systems. Artificial Intelligence, 195:496–527, 2013.

[Xiao et al., 2019] Tong Xiao, Yinqiao Li, Jingbo Zhu, Zhengtao Yu, and Tongran Liu. Sharing attention
  weights for fast transformer. In Proceedings of the Twenty-Eighth International Joint Conference on
  Artificial Intelligence (IJCAI-19), pages 5292–5298, 2019.
[Xie et al., 2022] Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma. An explanation
  of in-context learning as implicit bayesian inference. In Proceedings of International Conference on
  Learning Representations, 2022.
[Xin et al., 2020] Ji Xin, Raphael Tang, Jaejun Lee, Yaoliang Yu, and Jimmy Lin. Deebert: Dynamic early
  exiting for accelerating bert inference. In Proceedings of the 58th Annual Meeting of the Association for
  Computational Linguistics, pages 2246–2251, 2020.
[Xu et al., 2024] Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao,
  Qingwei Lin, and Daxin Jiang. Wizardlm: Empowering large pre-trained language models to follow
  complex instructions. In The Twelfth International Conference on Learning Representations, 2024.
[Yang et al., 2024] An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu,
  Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al. Qwen2. 5 technical report. arXiv preprint
  arXiv:2412.15115, 2024.
[Yang et al., 2019] Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and
  Quoc V Le. Xlnet: Generalized autoregressive pretraining for language understanding. Advances in
  neural information processing systems, 32, 2019.
[Yao et al., 2024] Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik
  Narasimhan. Tree of thoughts: Deliberate problem solving with large language models. Advances in
  Neural Information Processing Systems, 36, 2024.
[Yarowsky, 1995] David Yarowsky. Unsupervised word sense disambiguation rivaling supervised methods.
  In Proceedings of the 33rd annual meeting of the association for computational linguistics, pages 189–
  196, 1995.
[Yu et al., 2023] Zihan Yu, Liang He, Zhen Wu, Xinyu Dai, and Jiajun Chen. Towards better chain-of-
  thought prompting strategies: A survey. arXiv preprint arXiv:2310.04959, 2023.
[Zaheer et al., 2020] Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, C. Alberti,
  S. Ontañón, Philip Pham, Anirudh Ravula, Qifan Wang, L. Yang, and A. Ahmed. Big bird: Transformers
  for longer sequences. Advances in neural information processing systems, 33:17283–17297, 2020.
[Zellers et al., 2018] Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi. Swag: A large-scale
  adversarial dataset for grounded commonsense inference. In Proceedings of the 2018 Conference on
  Empirical Methods in Natural Language Processing, pages 93–104, 2018.
[Zhang and Sennrich, 2019] Biao Zhang and Rico Sennrich. Root mean square layer normalization. Ad-
  vances in Neural Information Processing Systems, 32, 2019.
[Zhang et al., 2023] Zhuosheng Zhang, Yao Yao, Aston Zhang, Xiangru Tang, Xinbei Ma, Zhiwei He,
  Yiming Wang, Mark Gerstein, Rui Wang, Gongshen Liu, and Hai Zhao. Igniting language intelli-
  gence: The hitchhiker’s guide from chain-of-thought reasoning to language agents. arXiv preprint
  arXiv:2311.11797, 2023a.
[Zhang et al., 2023] Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. Automatic chain of thought
  prompting in large language models. In The Eleventh International Conference on Learning Represen-
  tations, 2023b.
[Zhao et al., 2024] Hao Zhao, Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion. Long
  is more for alignment: A simple but tough-to-beat baseline for instruction fine-tuning. arXiv preprint
  arXiv:2402.04833, 2024.
[Zhao et al., 2023] Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou,
  Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Z. Chen,
  Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jianyun Nie, and Ji rong Wen.

  A survey of large language models. arXiv preprint arXiv:2303.18223, 2023.
[Zhou et al., 2023] Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma,
  Avia Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer
  Levy. Lima: Less is more for alignment. arXiv preprint arXiv:2305.11206, 2023a.
[Zhou et al., 2023] Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale
  Schuurmans, Claire Cui, Olivier Bousquet, Quoc V. Le, and Ed H. Chi. Least-to-most prompting enables
  complex reasoning in large language models. In Proceedings of The Eleventh International Conference
  on Learning Representations, 2023b.
[Zhou et al., 2020] Wangchunshu Zhou, Canwen Xu, Tao Ge, Julian McAuley, Ke Xu, and Furu Wei. Bert
  loses patience: Fast and robust inference with early exit. Advances in Neural Information Processing
  Systems, 33:18330–18341, 2020.
[Zhou et al., 2023] Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris
  Chan, and Jimmy Ba. Large language models are human-level prompt engineers. In The Eleventh
  International Conference on Learning Representations, 2023c.
[Zoph and Le, 2016] Barret Zoph and Quoc Le. Neural architecture search with reinforcement learning. In
  Proceedings of International Conference on Learning Representations, 2016.
[Zoph et al., 2020] Barret Zoph, Golnaz Ghiasi, Tsung-Yi Lin, Yin Cui, Hanxiao Liu, Ekin Dogus Cubuk,
  and Quoc Le. Rethinking pre-training and self-training. Advances in neural information processing
  systems, 33:3833–3845, 2020.

نمایهٔ کتاب (Index)

نمایهٔ پایانی کتاب و شمارهٔ صفحات ارجاع‌شده در نسخهٔ اصلی:

k-NN, 74                                few-shot COT prompting, 54
k-NN LM, 76
k-NN language modeling, 76              gated linear unit, 58
k-nearest neighbors, 74                 gaussian error linear unit, 58
                                        GeLU, 58
A2C, 175                                GLU, 58
action-value function, 171              GPT, 1
advantage, 175                          GQA, 80
advantage actor-critic, 175             Grouped query attention, 80
Agent, 47
                                        hard prompts, 140
ALiBi, 85
                                        human preference alignment, 152
alignment, 46
attention with linear biases, 85        ICL, 53
automated machine learning, 137         ICT, 6
automatic prompt design, 137            importance sampling, 180
AutoML, 137                             in-context learning, 6, 53, 95
autonomous agents, 134                  input inversion, 163
                                        instruction alignment, 152
BART, 19                                instruction fine-tuning, 43, 154
BERT, 1                                 interference, 30
Best-of-N sampling, 197                 internal memories, 74
BoN sampling, 197                       Interpolation, 82
Bradley-Terry model, 178                irreducible error, 63

calculation annotation, 113             key-value cache, 68
catastrophic forgetting, 34             KV cache, 68
causal language modeling, 9
chain of thought, 113                   label mapping, 105
chain-of-thought prompting, 53          Learning from Human Feedback, 47
completion, 6                           least-to-most prompting, 119
compositional generalization, 122       long-context LLMs, 66
CoT, 113                                masked language modeling, 1, 9
COT prompting, 53                       mBERT, 28
cross-lingual language models, 28       memory-based methods, 74
cumulative reward, 172                  MQA, 79
                                        multi-lingual BERT, 28
deliberate-then-generate, 126
                                        multi-query attention, 79
demonstrations, 6
direct preference optimization, 190     NAS, 137
Document Rotation, 20                   neural architecture search, 137
DPO, 190                                next sentence prediction, 13
DTG, 126                                NSP, 13

emergent abilities, 63                  offline reinforcement learning, 193
external memories, 74                   one-shot COT prompting, 54
Extrapolation, 81                       Outcome-based Approaches, 195

overoptimization problem, 189                 single-round prediction, 155
                                              soft prompts, 140
Performance Estimation, 137                   Span Masking, 19
performance function, 172                     state-value function, 171
performance gap recovered, 167                Strong Ceiling Performance, 167
permuted language modeling, 11                Sub-problem Generation, 118
PGR, 167                                      Sub-problem Solving, 118
Plackett-Luce model, 184                      superficial alignment hypothesis, 164
PPO, 50, 181                                  Supervised Fine-tuning, 47
prefix fine-tuning, 144                       supervised fine-tuning, 152
prefix language modeling, 16                  supervised learning, 2
problem decomposition, 116                    surrogate objective, 180
Process-based Approaches, 195
prompt embeddings, 148                        T5, 15
prompt engineering, 95                        TD, 176
prompt optimization, 137                      temporal difference, 176
Prompt Search Space, 137                      text completion, 109
prompting engineering, 51                     text transformation, 109
proximal policy optimization, 50, 181         Token Deletion, 19
                                              Token Masking, 19
Q-value function, 171                         Transformers, 1
                                              translation language modeling, 29
RAG, 76                                       trust regions, 181
ratio function, 180
rectified linear unit, 58                     unsupervised learning, 2
reinforcement learning from human feedback,
                                              Weak Performance, 167
          47, 153
                                              weak-to-strong generalization, 166
rejection sampling, 198
                                              Weak-to-strong Performance, 167
relation extraction, 108
ReLU, 58                                      XLMs, 28
retrieval-augmented generation, 76
return, 172                                   zero-shot COT, 54
reward gaming, 189                            zero-shot learning, 45
reward hacking, 189
Reward Model, 47
RLHF, 47, 153
RoBERTa, 26

sample efficient, 164
scaling laws, 63
self-consistency, 129
self-instruct, 160
self-supervised learning, 3
self-training, 3
Sentence Reordering, 19
Sequence Encoding Models, 3
Sequence Generation Models, 3
SFT, 47, 152

امتیاز کاربران به این مقاله

☆☆☆☆☆

0 نفر امتیاز داده اند. میانگین: 0.0 از 5

 

0 نظر

نظر محترم شما در مورد مقاله های وب سایت برنامه نویسی و پایگاه داده

نظرات محترم شما در خدمات رسانی بهتر ما را یاری می نمایند. لطفا اگر مایل بودید یک نظر ما را مهمان فرمائید. آدرس ایمیل و وب سایت شما نمایش داده نخواهد شد.

0 / 500

اطلاعات تماس

  • آدرس:اصفهان-خیابان ام کلثوم غربی - بعد خیابان تخم چی - بیست متر بعد از پیتزا ننه شب - کوچه تعمیر گاه سمار زغالی - پلاک 354 - درب مشکی - طبقه هفتم
  • آدرس ایمیل:najafzade@gmail.com
  • وب سایت:http://www.a00b.com/
  • تلفن ثابت:(+98)9131253620
  • تلفن همراه:09131253620