TOUVRON H, LAVRIL T, IZACARD G, et al. LLAma: Open and efficient foundation language models[J]. ArXiv Preprint ArXiv:2302, 13971, 2023.
WODAJO D, ATNAFU S, AKHTAR Z. Deepfake video detection using generative convolutional vision transformer[J].ArXiv Preprint ArXiv:2307, 07036, 2023.
DOSOVITSKIY A, BEYER L, KOLESNIKOV A, et al. An image is worth 16x16 words: Transformers for image recognition at scale[J]. ArXiv Preprint ArXiv: 2010. 11929, 2020.
DOSOVITSKIY A, BEYER L, KOLESNIKOV A, et al. An image is worth 16x16 words: Transformers for image recognition at scale[J]. ArXiv Preprint ArXiv:2010, 11929, 2020.