background image

Σχεδιασμός και Υλοποίηση εξειδικευμένου chatbot για διαδικτυακές πλατφόρμες

 

124

 

[48] 

Meta  AI.  (2024).  Llama  3.1  model  card:  υποστηριζόμενες  γλώσσες  [Διαδικτυακός  τόπος]. 

https://github.com/meta-llama/llama-models/blob/main/models/llama3_1/MODEL_CARD.md 

(πρόσβαση: Ιούνιος 2026).

 

[49] 

LLM Stats. (2026). Llama 3.1 8B Instruct vs Qwen2.5 7B Instruct: σύγκριση επιδόσεων (IFEval, 

MMLU-Pro  κ.ά.)  [Διαδικτυακός  τόπος].  https://llm-stats.com/models/compare/llama-3.1-8b-

instruct-vs-qwen-2.5-7b-instruct (πρόσβαση: Ιούνιος 2026).

 

[50] 

Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., & Steinhardt, J. (2021). 

Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300.

 

[51] 

Shieber,  S.  M.  (1994).  Lessons  from  a restricted  Turing  test.  Communications  of  the  ACM, 

37(6), 70–78.

 

[52] 

Spärck Jones, K. (1972). A statistical interpretation of term specificity and its application in 

retrieval. Journal of Documentation, 28(1), 11–21.

 

[53] 

Robertson, S. E., & Walker, S. (1994). Some simple effective approximations to the 2-Poisson 

model for probabilistic weighted retrieval. In Proceedings of the 17th Annual International ACM 

SIGIR Conference on Research and Development in Information Retrieval (pp. 232–241).

 

[54] 

Hoy, M. B. (2018). Alexa, Siri, Cortana, and more: An introduction to voice assistants. Medical 

Reference Services Quarterly, 37(1), 81–88.

 

[55] 

Kaplan,  J.,  et  al.  (2020).  Scaling  laws  for  neural  language  models.  arXiv  preprint 

arXiv:2001.08361.

 

[56] 

OpenAI. (2023). GPT-4 technical report. arXiv preprint arXiv:2303.08774.

 

[57] 

Touvron,  H.,  et  al.  (2023).  LLaMA:  Open  and  efficient  foundation  language  models.  arXiv 

preprint arXiv:2302.13971.

 

[58] 

Zhao, W. X., et al. (2023). A survey of large language models. arXiv preprint arXiv:2303.18223.

 

[59] 

Gholami, A., Kim, S., Dong, Z., Yao, Z., Mahoney, M. W., & Keutzer, K. (2021). A survey of 

quantization methods for efficient neural network inference. arXiv preprint arXiv:2103.13630.

 

[60] 

Frantar,  E.,  Ashkboos,  S., Hoefler,  T.,  &  Alistarh,  D.  (2022).  GPTQ:  Accurate post-training 

quantization for generative pre-trained transformers. arXiv preprint arXiv:2210.17323.

 

[61] 

Advanced  Micro  Devices,  Inc.  (2020).  “RDNA  2”  Instruction  Set  Architecture:  Reference 

Guide. Technical Report, AMD. https://docs.amd.com/v/u/en-US/rdna2-shader-instruction-set-

architecture (πρόσβαση: Ιούνιος 2026).

 

[62] 

Auffarth, B. (2023). Generative AI with LangChain. Birmingham, UK: Packt Publishing.

 

[63] 

Zirnstein,  B.  (2023).  Extended  context  for  InstructGPT  with  LlamaIndex.  Technical  Report. 

Hochschule für Wirtschaft und Recht Berlin.

 

[64] 

Kwon, W. (2025). vLLM: An efficient inference engine for large language models (Doctoral 

dissertation, UC Berkeley).

 

[65] 

Academic  Reference.  (2026).  Which  quantization  should  I  use?  A  unified  evaluation  of 

llama.cpp quantization on Llama-3.1-8B-Instruct. arXiv preprint arXiv:2601.14277.