[ { "text": "FaceForensics++: Learning to Detect Manipulated Facial Images Andreas R\u00a8ossler1 Davide Cozzolino2 Luisa Verdoliva2 Christian Riess3 Justus Thies1 Matthias Nie\u00dfner1 1Technical University of Munich 2University Federico II of Naples 3University of Erlangen-Nuremberg Figure 1: FaceForensics++ is a dataset of facial forgeries that enables researchers to train deep-learning-based approaches", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 0 }, { "text": "in a supervised fashion. The dataset contains manipulations created with four state-of-the-art methods, namely, Face2Face, FaceSwap, DeepFakes, and NeuralTextures. Abstract The rapid progress in synthetic image generation and manipulation has now come to a point where it raises signif- icant concerns for the implications towards society. At best,", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 1 }, { "text": "this leads to a loss of trust in digital content, but could po- tentially cause further harm by spreading false information or fake news. This paper examines the realism of state-of- the-art image manipulations, and how dif\ufb01cult it is to detect them, either automatically or by humans. To standardize the evaluation of detection methods, we", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 2 }, { "text": "propose an automated benchmark for facial manipulation detection1. In particular, the benchmark is based on Deep- Fakes [1], Face2Face [59], FaceSwap [2] and NeuralTex- tures [57] as prominent representatives for facial manipula- tions at random compression level and size. The benchmark is publicly available2 and contains a hidden test set as well", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 3 }, { "text": "as a database of over 1.8 million manipulated images. This dataset is over an order of magnitude larger than compara- ble, publicly available, forgery datasets. Based on this data, we performed a thorough analysis of data-driven forgery detectors. We show that the use of additional domain- speci\ufb01c knowledge improves forgery detection to unprece-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 4 }, { "text": "dented accuracy, even in the presence of strong compres- sion, and clearly outperforms human observers. 1. Introduction Manipulation of visual content has now become ubiqui- tous, and one of the most critical topics in our digital so- ciety. For instance, DeepFakes [1] has shown how com- puter graphics and visualization techniques can be used to", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 5 }, { "text": "defame persons by replacing their face by the face of a dif- ferent person. Faces are of special interest to current manip- ulation methods for various reasons: \ufb01rstly, the reconstruc- tion and tracking of human faces is a well-examined \ufb01eld in computer vision [68], which is the foundation of these", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 6 }, { "text": "editing approaches. Secondly, faces play a central role in human communication, as the face of a person can empha- size a message or it can even convey a message in its own right [28]. Current facial manipulation methods can be separated into two categories: facial expression manipulation and fa- cial identity manipulation (see Fig. 2). One of the most", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 7 }, { "text": "prominent facial expression manipulation techniques is the method of Thies et al. [59] called Face2Face. It enables the transfer of facial expressions of one person to another per- son in real time using only commodity hardware. Follow-up work such as \u201cSynthesizing Obama\u201d [55] is able to animate the face of a person based on an audio input sequence.", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 8 }, { "text": "prominent and widespread identity editing technique is face swapping, which has gained signi\ufb01cant popularity as lightweight systems are now capable of running on mobile phones. Additionally, facial reenactment techniques are now available, which alter the expressions of a person by transferring the expressions of a source person to the target.", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 9 }, { "text": "Identity manipulation is the second category of facial forgeries. Instead of changing expressions, these methods replace the face of a person with the face of another per- son. This category is known as face swapping. It became popular with wide-spread consumer-level applications like Snapchat. DeepFakes also performs face swapping, but via", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 10 }, { "text": "deep learning. While face swapping based on simple com- puter graphics techniques can run in real time, DeepFakes need to be trained for each pair of videos, which is a time- consuming task. In this work, we show that we can automatically and re- liably detect such manipulations, and thereby outperform", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 11 }, { "text": "human observers by a signi\ufb01cant margin. We leverage re- cent advances in deep learning, in particular, the ability to learn extremely powerful image features with convolutional neural networks (CNNs). We tackle the detection problem by training a neural network in a supervised fashion. To this end, we generate a large-scale dataset of manipulations", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 12 }, { "text": "based on the classical computer graphics-based methods Face2Face [59] and FaceSwap [2] as well as the learning- based approaches DeepFakes [1] and NeuralTextures [57]. As the digital media forensics \ufb01eld lacks a benchmark for forgery detection, we propose an automated benchmark that considers the four manipulation methods in a realistic", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 13 }, { "text": "scenario, i.e., with random compression and random dimensions. Using this benchmark, we evaluate the current state-of-the-art detection methods as well as our forgery detection pipeline that considers the restricted \ufb01eld of facial manipulation methods. Our paper makes the following contributions: \u2022 an automated benchmark for facial manipulation de-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 14 }, { "text": "tection under random compression for a standardized comparison, including a human baseline, \u2022 a novel large-scale dataset of manipulated facial im- agery composed of more than 1.8 million images from 1,000 videos with pristine (i.e., real) sources and tar- get ground truth to enable supervised learning,", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 15 }, { "text": "\u2022 an extensive evaluation of state-of-the-art hand-crafted and learned forgery detectors in various scenarios, \u2022 a state-of-the-art forgery detection method tailored to facial manipulations. 2. Related Work The paper intersects several \ufb01elds in computer vision and digital multimedia forensics. We cover the most important", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 16 }, { "text": "related papers in the following paragraphs. Face Manipulation Methods: In the last two decades, in- terest in virtual face manipulation has rapidly increased. A comprehensive state-of-the-art report has been published by Zollh\u00a8ofer et al. [68]. In particular, Bregler et al. [13] pre- sented an image-based approach called Video Rewrite to au-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 17 }, { "text": "tomatically create a new video of a person with generated mouth movements. With Video Face Replacement [20], Dale et al. presented one of the \ufb01rst automatic face swap methods. Using single-camera videos, they reconstruct a 3D model of both faces and exploit the corresponding 3D geometry to warp the source face to the target face. Gar-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 18 }, { "text": "rido et al. [29] presented a similar system that replaces the face of an actor while preserving the original expressions. VDub [30] uses high-quality 3D face capturing techniques to photo-realistically alter the face of an actor to match the mouth movements of a dubber. Thies et al. [58] demon- strated the \ufb01rst real-time expression transfer for facial reen-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 19 }, { "text": "actment. Based on a consumer level RGB-D camera, they reconstruct and track a 3D model of the source and the target actor. The tracked deformations of the source face are applied to the target face model. As a \ufb01nal step, they blend the altered face on top of the original target video. Face2Face, proposed by Thies et al. [59], is an advanced", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 20 }, { "text": "real-time facial reenactment system, capable of altering fa- cial movements in commodity video streams, e.g., videos from the internet. They combine 3D model reconstruction and image-based rendering techniques to generate their out- put. The same principle can be also applied in Virtual Real- ity in combination with eye-tracking and reenactment [60]", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 21 }, { "text": "or be extended to the full body [61]. Kim et al. [39] learn an image-to-image translation network to convert computer graphic renderings of faces to real images. Instead of a pure image-to-image translation network, NeuralTextures [57] optimizes a neural texture in conjunction with a rendering network to compute the reenactment result. In compari-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 22 }, { "text": "son to Deep Video Portraits [39], it shows sharper results, especially, in the mouth region. Suwajanakorn et al. [55] learned the mapping between audio and lip motions, while their compositing approach builds on similar techniques to Face2Face [59]. Averbuch-Elor et al. [8] present a reen- actment method, Bringing Portraits to Life, which employs", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 23 }, { "text": "2D warps to deform the image to match the expressions of a source actor. They also compare to the Face2Face technique and achieve similar quality. Recently, several face image synthesis approaches us- ing deep learning techniques have been proposed. Lu et al. [47] provide an overview. Generative adversarial net-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 24 }, { "text": "works (GANs) are used to apply Face Aging [7], to gener- ate new viewpoints [34], or to alter face attributes like skin color [46]. Deep Feature Interpolation [62] shows impres- sive results on altering face attributes like age, mustache, smiling etc. Similar results of attribute interpolations are", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 25 }, { "text": "achieved by Fader Networks [43]. Most of these deep learn- ing based image synthesis techniques suffer from low image resolutions. Recently, Karras et al. [37] have improved the image quality using progressive growing of GANs, produc- ing high-quality synthesis of faces. Multimedia Forensics: Multimedia forensics aims to en-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 26 }, { "text": "sure authenticity, origin, and provenance of an image or video without the help of an embedded security scheme. Focusing on integrity, early methods are driven by hand- crafted features that capture expected statistical or physics- based artifacts that occur during image formation. Surveys on these methods can be found in [26, 53]. More recent lit-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 27 }, { "text": "erature concentrates on CNN-based solutions, through both supervised and unsupervised learning [10, 17, 12, 9, 35, 67]. For videos, the main body of work focuses on detecting ma- nipulations that can be created with relatively low effort, such as dropped or duplicated frames [63, 31, 45], varying interpolation types [25], copy-move manipulations [11, 21],", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 28 }, { "text": "or chroma-key compositions [48]. Several other works explicitly refer to detecting manip- ulations related to faces, such as distinguishing computer generated faces from natural ones [22, 15, 51], morphed faces [50], face splicing [24, 23], face swapping [66, 38] and DeepFakes [5, 44, 33]. For face manipulation detec-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 29 }, { "text": "tion, some approaches exploit speci\ufb01c artifacts arising from the synthesis process, such as eye blinking [44], or color, texture and shape cues [24, 23]. Other works are more gen- eral and propose a deep network trained to capture the sub- tle inconsistencies arising from low-level and/or high level", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 30 }, { "text": "features [50, 66, 38, 5, 33]. These approaches show im- pressive results, however robustness issues often remain un- addressed, although they are of paramount importance for practical applications. For example, operations like com- pression and resizing are known for laundering manipula- tion traces from the data. In real-world scenarios, these", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 31 }, { "text": "basic operations are standard when images and videos are for example uploaded to social media, which is one of the most important application \ufb01eld for forensic analysis. To this end, our dataset is designed to cover such realistic sce- narios, i.e., videos from the wild, manipulated and com- pressed with different quality levels (see Section 3). The", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 32 }, { "text": "availability of such a large and varied dataset can help re- searchers to benchmark their approaches and develop better forgery detectors for facial imagery. Forensic Analysis Datasets: Classical forensics datasets have been created with signi\ufb01cant manual effort under very controlled conditions, to isolate speci\ufb01c properties of the", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 33 }, { "text": "data like camera artifacts. While several datasets were proposed that include image manipulations, only a few of them also address the important case of video footage. MICC F2000, for example, is an image copy-move manip- ulation dataset consisting of a collection of 700 forged im- ages from various sources [6]. The First IEEE Image Foren-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 34 }, { "text": "sics Challenge Dataset comprises a total of 1176 forged images; the Wild Web Dataset [64] with 90 real cases of manipulations coming from the web and the Realistic Tampering dataset [42] including 220 forged images. A database of 2010 FaceSwap- and SwapMe-generated im- ages has been proposed by Zhou et al. [66]. Recently, Kor-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 35 }, { "text": "shunov and Marcel [41] constructed a dataset of 620 Deep- fakes videos created from multiple videos for each of 43 subjects. The National Institute of Standards and Technol- ogy (NIST) released the most extensive dataset for generic image manipulation comprising about 50, 000 forged im- ages (both local and global manipulations) and around 500", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 36 }, { "text": "forged videos [32]. In contrast, we construct a database containing more than 1.8 million images from 4000 fake videos \u2013 an order of magnitude more than existing datasets. We evaluate the im- portance of such a large training corpus in Section 4. 3. Large-Scale Facial Forgery Database A core contribution of this paper is our FaceForensics++", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 37 }, { "text": "(a) Gender (b) Resolution (c) Pixel Coverage of Faces Figure 3: Statistics of our sequences. VGA denotes 480p, HD denotes 720p, and FHD denotes 1080p resolution of our videos. The graph (c) shows the number of sequences (y-axis) with given bounding box pixel height (x-axis). [52]. This new large-scale dataset enables us to train a state-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 38 }, { "text": "of-the-art forgery detector for facial image manipulation in a supervised fashion (see Section 4). To this end, we lever- age four automated state-of-the-art face manipulation meth- ods, which are applied to 1,000 pristine videos downloaded from the Internet (see Fig. 3 for statistics). To imitate realis-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 39 }, { "text": "tic scenarios, we chose to collect videos in the wild, specif- ically from YouTube. However, early experiments with all manipulation methods showed that the target face had to be nearly front-facing to prevent the manipulation methods from failing or producing strong artifacts. Thus, we per- form a manual screening of the resulting clips to ensure a", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 40 }, { "text": "high-quality video selection and to avoid videos with face occlusions. We selected 1,000 video sequences containing 509, 914 images which we use as our pristine data. To generate a large scale manipulation database, we adapted state-of-the-art video editing methods to work fully automatically. In the following paragraphs, we brie\ufb02y de-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 41 }, { "text": "scribe these methods. For our dataset, we chose two computer graphics-based approaches (Face2Face and FaceSwap) and two learning- based approaches (DeepFakes and NeuralTextures). All four methods require source and target actor video pairs as input. The \ufb01nal output of each method is a video composed", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 42 }, { "text": "of generated images. Besides the manipulation output, we also compute ground truth masks that indicate whether a pixel has been modi\ufb01ed or not, which can be used to train forgery localization methods. For more information and hyper-parameters we refer to Appendix D. FaceSwap FaceSwap is a graphics-based approach to", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 43 }, { "text": "transfer the face region from a source video to a target video. Based on sparse detected facial landmarks the face region is extracted. Using these landmarks, the method \ufb01ts a 3D template model using blendshapes. This model is back- projected to the target image by minimizing the difference between the projected shape and the localized landmarks", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 44 }, { "text": "using the textures of the input image. Finally, the rendered model is blended with the image and color correction is ap- plied. We perform these steps for all pairs of source and target frames until one video ends. The implementation is computationally lightweight and can be ef\ufb01ciently run on the CPU.", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 45 }, { "text": "DeepFakes The term Deepfakes has widely become a synonym for face replacement based on deep learning, but it is also the name of a speci\ufb01c manipulation method that was spread via online forums. To distinguish these, we denote said method by DeepFakes in the following paper. There are various public implementations of DeepFakes", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 46 }, { "text": "available, most notably FakeApp [3] and the faceswap github [1]. A face in a target sequence is replaced by a face that has been observed in a source video or image col- lection. The method is based on two autoencoders with a shared encoder that are trained to reconstruct training im- ages of the source and the target face, respectively. A face", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 47 }, { "text": "detector is used to crop and to align the images. To create a fake image, the trained encoder and decoder of the source face are applied to the target face. The autoencoder output is then blended with the rest of the image using Poisson image editing [49]. For our dataset, we use the faceswap github implemen-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 48 }, { "text": "tation. We slightly modify the implementation by replacing the manual training data selection with a fully automated data loader. We used the default parameters to train the video-pair models. Since the training of these models is very time-consuming, we also publish the models as part of the dataset. This facilitates generation of additional manip-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 49 }, { "text": "ulations of these persons with different post-processing. Face2Face Face2Face [59] is a facial reenactment system that transfers the expressions of a source video to a target video while maintaining the identity of the target person. The original implementation is based on two video input streams, with manual keyframe selection. These frames are", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 50 }, { "text": "used to generate a dense reconstruction of the face which can be used to re-synthesize the face under different illumi- nation and expressions. To process our video database, we adapt the Face2Face approach to fully-automatically cre- ate reenactment manipulations. We process each video in a preprocessing pass; here, we use the \ufb01rst frames in order", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 51 }, { "text": "to obtain a temporary face identity (i.e., a 3D model), and track the expressions over the remaining frames. In order to select the keyframes required by the approach, we automat- ically select the frames with the left- and right-most angle of the face. Based on this identity reconstruction, we track", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 52 }, { "text": "the whole video to compute per frame the expression, rigid pose, and lighting parameters as done in the original im- plementation of Face2Face. We generate the reenactment video outputs by transferring the source expression param- eters of each frame (i.e., 76 Blendshape coef\ufb01cients) to the target video. More details of the reenactment process can", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 53 }, { "text": "NeuralTextures Thies et al. [57] show facial reenactment as an example for their NeuralTextures-based rendering ap- proach. It uses the original video data to learn a neu- ral texture of the target person, including a rendering net- work. This is trained with a photometric reconstruction loss in combination with an adversarial loss. In our im-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 54 }, { "text": "plementation, we apply a patch-based GAN-loss as used in Pix2Pix [36]. The NeuralTextures approach relies on tracked geometry that is used during train and test times. We use the tracking module of Face2Face to generate these information. We only modify the facial expressions corre- sponding to the mouth region, i.e., the eye region stays un-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 55 }, { "text": "changed (otherwise the rendering network would need con- ditional input for the eye movement similar to Deep Video Portraits [39]). Postprocessing - Video Quality To create a realistic set- ting for manipulated videos, we generate output videos with different quality levels, similar to the video processing of", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 56 }, { "text": "many social networks. Since raw videos are rarely found on the internet, we compress the videos using the H.264 codec, which is widely used by social networks or video- sharing websites. To generate high quality videos, we use a light compression denoted by HQ (constant rate quantiza- tion parameter equal to 23) which is visually nearly lossless.", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 57 }, { "text": "Low quality videos (LQ ) are produced using a quantization of 40. 4. Forgery Detection We cast the forgery detection as a per-frame binary clas- si\ufb01cation problem of the manipulated videos. The following sections show the results of manual and automatic forgery detection. For all experiments, we split the dataset into a", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 58 }, { "text": "\ufb01xed training, validation, and test set, consisting of 720, 140, and 140 videos respectively. All evaluations are re- ported using videos from the test set. For all graphs, we list the exact numbers in Appendix B. 4.1. Forgery Detection of Human Observers To evaluate the performance of humans in the task of", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 59 }, { "text": "forgery detection, we conducted a user study with 204 par- ticipants consisting mostly of computer science university students. This forms the baseline for the automated forgery detection methods. Layout of the User Study: After a short introduction to the binary task, users are instructed to classify randomly se-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 60 }, { "text": "lected images from our test set. The selected images vary in image quality as well as manipulation method; we used a 50:50 split of pristine and fake images. Since the amount time for inspection of an image may be important, and to mimic scenario where a user only spends a limited amount of time per image as is common on social media, we ran-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 61 }, { "text": "domly set a time limit of 2, 4 or 6 seconds after which we hide the image. Afterwards, the users were asked whether the displayed image is \u2018real\u2019 or \u2018fake\u2019. To ensure that the users spend the available time on inspection, the question is asked after the image has been displayed and not during the observation time. We designed the study to only take a few", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 62 }, { "text": "minutes per participant, showing 60 images per attendee, which results in a collection of 12240 human decisions. Evaluation: In Fig. 4, we show the results of our study on all quality levels, showing a correlation between video quality and the ability to detect fakes. With a lower video quality, the human performance decreases in average from", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 63 }, { "text": "68.7% to 58.7%. The graph shows the numbers averaged across all time intervals since the different time constraints did not result in signi\ufb01cantly different observations. Figure 4: Forgery detection results of our user study with 204 participants. The accuracy is dependent on the video quality and results in a decreasing accuracy rate that is", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 64 }, { "text": "68.69% in average on raw videos, 66.57% on high quality, and 58.73% on low quality videos. Note that the user study contained fake images of all four manipulation methods and pristine images. In this setting, Face2Face and NeuralTextures were particularly dif\ufb01cult to detect by human observers, as they do not introduce a strong", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 65 }, { "text": "semantic change, introducing only subtle visual artifacts in contrast to the face replacement methods. NeuralTextures texture seems particularly dif\ufb01cult to detect as human de- tection accuracy is below random chance and only increases in the challenging low quality task. 4.2. Automatic Forgery Detection Methods", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 66 }, { "text": "Our forgery detection pipeline is depicted in Fig. 5. Since our goal is to detect forgeries of facial imagery, we use additional domain-speci\ufb01c information that we can ex- tract from input sequences. To this end, we use the state- of-the-art face tracking method by Thies et al. [59] to track the face in the video and to extract the face region of the", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 67 }, { "text": "improves the overall performance of a forgery detector in comparison to a na\u00a8\u0131ve approach that uses the whole image as input (see Sec. 4.2.2). We evaluated various variants of our approach by using different state-of-the-art classi\ufb01ca- tion methods. We are considering learning-based methods used in the forensic community for generic manipulation", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 68 }, { "text": "detection [10, 17], computer-generated vs natural image de- tection [51] and face tampering detection [5]. In addition, we show that the classi\ufb01cation based on XceptionNet [14] outperforms all other variants in detecting fakes. 4.2.1 Detection based on Steganalysis Features: We evaluate detection from steganalysis features, follow-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 69 }, { "text": "ing the method by Fridrich et al. [27] which employs hand- crafted features. The features are co-occurrences on 4 pixels patterns along the horizontal and vertical direction on the high-pass images for a total feature length of 162. These features are then used to train a linear Support Vector Ma- chine (SVM) classi\ufb01er. This technique was the winning", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 70 }, { "text": "approach in the \ufb01rst IEEE Image Forensic Challenge [16]. We provide a 128 \u00d7 128 central crop-out of the face as in- put to the method. While the hand-crafted method outper- forms human accuracy on raw images by a large margin, it struggles to cope with compression, which leads to an accu- racy below human performance for low quality videos (see", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 71 }, { "text": "Fig. 6 and Table 1). 4.2.2 Detection based on Learned Features: For detection from learned features, we evaluate \ufb01ve net- work architectures known from the literature to solve the classi\ufb01cation task: (1) Cozzolino et al. [17] cast the hand-crafted Steganal- ysis features from the previous section to a CNN-based net-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 72 }, { "text": "work. We \ufb01ne-tune this network on our large scale dataset. (2) We use our dataset to train the convolutional neu- ral network proposed by Bayar and Stamm [10] that uses a constrained convolutional layer followed by two convo- lutional, two max-pooling and three fully-connected layers. The constrained convolutional layer is speci\ufb01cally designed", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 73 }, { "text": "See Table 1 for the average accuracy values. Aside from the Full Image XceptionNet, we use the proposed pre- extraction of the face region as input to the approaches. to suppress the high-level content of the image. Similar to the previous methods, we use a centered 128 \u00d7 128 crop as input. (3) Rahmouni et al. [51] adopt different CNN architec-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 74 }, { "text": "tures with a global pooling layer that computes four statis- tics (mean, variance, maximum and minimum). We con- sider the Stats-2L network that had the best performance. (4) MesoInception-4 [5] is a CNN-based network in- spired by InceptionNet [56] to detect face tampering in videos. The network has two inception modules and two", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 75 }, { "text": "stead of the classic cross-entropy loss, the authors propose the mean squared error between true and predicted labels. We resize the face images to 256 \u00d7 256, the input of the network. (5) XceptionNet [14] is a traditional CNN trained on Im- ageNet based on separable convolutions with residual con-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 76 }, { "text": "nections. We transfer it to our task by replacing the \ufb01nal fully connected layer with two outputs. The other layers are initialized with the ImageNet weights. To set up the newly inserted fully connected layer, we \ufb01x all weights up to the \ufb01- nal layers and pre-train the network for 3 epochs. After this", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 77 }, { "text": "step, we train the network for 15 more epochs and choose the best performing model based on validation accuracy. A detailed description of our training and hyper- parameters can be found in Appendix D. Comparison of our Forgery Detection Variants: Fig. 6 shows the results of a binary forgery detection task using", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 78 }, { "text": "all network architectures evaluated separately on all four manipulation methods and at different video quality levels. All approaches achieve very high performance on raw input data. Performance drops for compressed videos, particu- larly for hand-crafted features and for shallow CNN archi- tectures [10, 17]. The neural networks are better at han-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 79 }, { "text": "dling these situations, with XceptionNet able to achieve compelling results on weak compression while still main- taining reasonable performance on low quality images, as it bene\ufb01ts from its pre-training on ImageNet as well as larger network capacity. To compare the results of our user study to the perfor-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 80 }, { "text": "mance of our automatic detectors, we also tested the detec- tion variants on a dataset containing images from all ma- nipulation methods. Fig. 7 and Table 1 show the results on the full dataset. Here, our automated detectors outperform human performance by a large margin (cf. Fig. 4). We also evaluate a na\u00a8\u0131ve forgery detector operating on the full im-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 81 }, { "text": "age (resized to the XceptionNet input) instead of using face tracking information (see Fig. 7, rightmost column). Due to the lack of domain-speci\ufb01c information, the XceptionNet classi\ufb01er has a signi\ufb01cantly lower accuracy in this scenario. To summarize, domain-speci\ufb01c information in combination with a XceptionNet classi\ufb01er shows the best performance", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 82 }, { "text": "in each test. We use this network to further understand the in\ufb02uence of the training corpus size and its ability to distin- guish between the different manipulation methods. Forgery Detection of GAN-based methods The experi- ments show that all detection approaches achieve a lower accuracy on the GAN-based NeuralTextures approach.", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 83 }, { "text": "NeuralTextures is training a unique model for every ma- nipulation which results in a higher variation of possible artifacts. While DeepFakes is also training one model per manipulation, it uses a \ufb01xed post-processing pipeline sim- Compression Raw HQ LQ [14] XceptionNet Full Image 82.01 74.78 70.52", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 84 }, { "text": "full image XceptionNet, all methods are trained on a con- servative crop (enlarged by a factor of 1.3) around the center of the tracked face. Figure 8: The detection performance of our approach us- ing XceptionNet depends on the training corpus size. Espe- cially, for low quality video data, a large database is needed.", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 85 }, { "text": "ilar to the computer-based manipulation methods and thus has consistent artifacts. Evaluation of the Training Corpus Size: Fig. 8 shows the importance of the training corpus size. To this end, we trained the XceptionNet classi\ufb01er with different train- ing corpus sizes on all three video quality level separately.", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 86 }, { "text": "5. Benchmark In addition to our large-scale manipulation database, we publish a competitive benchmark for facial forgery detec- tion. To this end, we collected 1000 additional videos and manipulated a subset of those in a similar fashion as in Section 3 for each of our four manipulation methods. As uploaded videos (e.g., to social networks) will be post-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 87 }, { "text": "processed in various ways, we obscure all selected videos multiple times (e.g., by unknown re-sizing, compression method and bit-rate) to ensure realistic conditions. This processing is directly applied on raw videos. Finally, we manually select a single challenging frame from each video based on visual inspection. Speci\ufb01cally, we collect a set of", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 88 }, { "text": "1000 images, each image randomly taken from either the manipulation methods or the original footage. Note that we do not necessarily have an equal split of pristine and fake images nor an equal split of the used manipulation meth- ods. The ground truth labels are hidden and are used on our host server to evaluate the classi\ufb01cation accuracy of the", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 89 }, { "text": "submitted models. The automated benchmark allows sub- missions every two weeks from a single submitter to prevent over\ufb01tting (similar to existing benchmarks [19]). As baselines, we evaluate the low quality versions of our previously trained models on the benchmark and report the numbers for each detection method separately (see Ta-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 90 }, { "text": "ble 2). Aside from the Full Image XceptionNet, we use the proposed pre-extraction of the face region as input to the approaches. The relative performance of the classi\ufb01ca- tion models is similar to our database test set (see Table 1). However, since the benchmark scenario deviates from the training database, the overall performance of the models", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 91 }, { "text": "is lower, especially for the pristine image detection preci- sion; the major changes being the randomized quality level as well as possible tracking errors during test. Since our proposed method relies on face detections, we predict fake as default in case of a tracking failure. The benchmark is already publicly available to the com-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 92 }, { "text": "munity and we hope that it leads to a standardized compar- ison of follow-up work. 6. Discussion & Conclusion While current state-of-the-art facial image manipulation methods exhibit visually stunning results, we demonstrate that they can be detected by trained forgery detectors. It is particularly encouraging that also the challenging case", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 93 }, { "text": "of low-quality video can be tackled by learning-based ap- proaches, where humans and hand-crafted features exhibit dif\ufb01culties. To train detectors using domain-speci\ufb01c knowl- edge, we introduce a novel dataset of videos of manipulated faces that exceeds all existing publicly available forensic datasets by an order of magnitude.", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 94 }, { "text": "In this paper, we focus on the in\ufb02uence of compression to Accuracies DF F2F FS NT Real Total Xcept. Full Image 74.55 75.91 70.87 73.33 51.00 62.40 Steg. Features 73.64 73.72 68.93 63.33 34.00 51.80 Cozzolino et al. 85.45 67.88 73.79 78.00 34.40 55.20 Rahmouni et al. 85.45 64.23 56.31 60.07 50.00 58.10", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 95 }, { "text": "Bayar and Stamm 84.55 73.72 82.52 70.67 46.20 61.60 MesoNet 87.27 56.20 61.17 40.67 72.60 66.00 XceptionNet 96.36 86.86 90.29 80.67 52.40 70.10 Table 2: Results of the low quality trained model of each detection method on our benchmark. We report precision results for DeepFakes (DF), Face2Face (F2F), FaceSwap", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 96 }, { "text": "(FS), NeuralTextures (NT), and pristine images (Real) as well as the overall total accuracy. the detectability of state-of-the-art manipulation methods, proposing a standardized benchmark for follow-up work. All image data, trained models, as well as our benchmark are publicly available and are already used by other re-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 97 }, { "text": "searchers. In particular, transfer learning is of high inter- est in the forensic community. As new manipulation meth- ods appear by the day, methods must be developed that are able to detect fakes with little to no training data. Our database is already used for this forensic transfer learning task, where knowledge of one source manipulation domain", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 98 }, { "text": "is transferred to another target domain, as shown by Coz- zolino et al [18]. We hope that the dataset and benchmark become a stepping stone for future research in the \ufb01eld of digital media forensics, and in particular with a focus on facial forgeries. 7. Acknowledgement We gratefully acknowledge the support of this research", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 99 }, { "text": "by the AI Foundation, a TUM-IAS Rudolf M\u00a8o\u00dfbauer Fel- lowship, the ERC Starting Grant Scan2CAD (804724), and a Google Faculty Award. We would also like to thank Google\u2019s Chris Bregler for help with the cloud computing. In addition, this material is based on research sponsored by the Air Force Research Laboratory and the Defense Ad-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 100 }, { "text": "vanced Research Projects Agency under agreement num- ber FA8750-16-2-0204. The U.S. Government is authorized to reproduce and distribute reprints for Governmental pur- poses notwithstanding any copyright notation thereon. The views and conclusions contained herein are those of the au- thors and should not be interpreted as necessarily represent-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 101 }, { "text": "Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan. YouTube-8m: A large- scale video classi\ufb01cation benchmark. arXiv preprint arXiv:1609.08675, 2016. 12 [5] Darius Afchar, Vincent Nozick, Junichi Yamagishi, and Isao Echizen. Mesonet: a compact facial video forgery detection", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 102 }, { "text": "network. arXiv preprint arXiv:1809.00888, 2018. 3, 6, 7, 13, 14 [6] Irene Amerini, Lamberto Ballan, Roberto Caldelli, Alberto Del Bimbo, and Giuseppe Serra. A SIFT-based forensic method for copy-move attack detection and transformation recovery. IEEE Transactions on Information Forensics and Security, 6(3):1099\u20131110, Mar. 2011. 3", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 103 }, { "text": "tions on Graphics (Proceeding of SIGGRAPH Asia 2017), 36(4):to appear, 2017. 3 [9] Jawadul H. Bappy, Amit K. Roy-Chowdhury, Jason Bunk, Lakshmanan Nataraj, and B.S. Manjunath. Exploiting spatial structure for localizing manipulated image regions. In IEEE International Conference on Computer Vision, pages 4970\u2013", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 104 }, { "text": "4979, 2017. 3 [10] Belhassen Bayar and Matthew C. Stamm. A deep learning approach to universal image manipulation detection using a new convolutional layer. In ACM Workshop on Information Hiding and Multimedia Security, pages 5\u201310, 2016. 3, 6, 7, 13, 14 [11] Paolo Bestagini, Simone Milani, Marco Tagliasacchi, and", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 105 }, { "text": "Stefano Tubaro. Local tampering detection in video se- quences. In IEEE International Workshop on Multimedia Signal Processing, pages 488\u2013493, October 2013. 3 [12] Luca Bondi, Silvia Lameri, David G\u00a8uera, Paolo Bestagini, Edward J. Delp, and Stefano Tubaro. Tampering Detection and Localization through Clustering of Camera-Based CNN", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 106 }, { "text": "Features. In IEEE Computer Vision and Pattern Recognition Workshops, 2017. 3 [13] Christoph Bregler, Michele Covell, and Malcolm Slaney. Video rewrite: Driving visual speech with audio. In 24th Annual Conference on Computer Graphics and Interactive Techniques, SIGGRAPH \u201997, pages 353\u2013360, 1997. 2 [14] Francois Chollet. Xception: Deep Learning with Depthwise", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 107 }, { "text": "Separable Convolutions. In IEEE Conference on Computer Vision and Pattern Recognition, 2017. 6, 7, 13, 14 [15] Valentina Conotter, Ecaterina Bodnari, Giulia Boato, and Hany Farid. Physiologically-based detection of computer generated faces in video. In IEEE International Conference on Image Processing, pages 1\u20135, Oct 2014. 3", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 108 }, { "text": "Recasting residual-based local descriptors as convolutional neural networks: an application to image forgery detection. In ACM Workshop on Information Hiding and Multimedia Security, pages 1\u20136, 2017. 3, 6, 7, 13, 14 [18] Davide Cozzolino, Justus Thies, Andreas R\u00a8ossler, Chris- tian Riess, Matthias Nie\u00dfner, and Luisa Verdoliva. Foren-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 109 }, { "text": "sicTransfer: Weakly-supervised Domain Adaptation for Forgery Detection. arXiv preprint arXiv:1812.02510, 2018. 8 [19] Angela Dai, Angel X. Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nie\u00dfner. ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes. In IEEE Computer Vision and Pattern Recognition, 2017. 8", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 110 }, { "text": "IEEE Transactions on Circuits and Systems for Video Tech- nology, in press, 2018. 3 [22] Duc-Tien Dang-Nguyen, Giulia Boato, and Francesco De Natale. Identify computer generated characters by analysing facial expressions variation. In IEEE International Work- shop on Information Forensics and Security, pages 252\u2013257,", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 111 }, { "text": "2012. 3 [23] Tiago de Carvalho, Fabio A. Faria, Helio Pedrini, Ricardo da S. Torres, and Anderson Rocha. Illuminant-Based Trans- formed Spaces for Image Forensics. IEEE Transactions on Information Forensics and Security, 11(4):720\u2013733, 2016. 3 [24] Tiago de Carvalho, Christian Riess, Elli Angelopoulou, He-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 112 }, { "text": "lio Pedrini, and Anderson Rocha. Exposing digital image forgeries by illumination color classi\ufb01cation. IEEE Trans- actions on Information Forensics and Security, 8(7):1182\u2013 1194, 2013. 3 [25] Xiangling Ding, Gaobo Yang, Ran Li, Lebing Zhang, Yue Li, and Xingming Sun. Identi\ufb01cation of Motion-Compensated", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 113 }, { "text": "Frame Rate Up-Conversion Based on Residual Signal. IEEE Transactions on Circuits and Systems for Video Technology, in press, 2017. 3 [26] Hany Farid. Photo Forensics. The MIT Press, 2016. 3 [27] Jessica Fridrich and Jan Kodovsk\u00b4y. Rich Models for Ste- ganalysis of Digital Images. IEEE Transactions on Informa-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 114 }, { "text": "and Pattern Recognition, pages 4217\u20134224, 2014. 2 [30] Pablo Garrido, Levi Valgaerts, Hamid Sarmadi, Ingmar Steiner, Kiran Varanasi, Patrick P\u00b4erez, and Christian Theobalt. Vdub: Modifying face video of actors for plau- sible visual alignment to a dubbed audio track. Computer Graphics Forum, 34(2):193\u2013204, 2015. 2", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 115 }, { "text": "Amy N. Yates, Andrew Delgado, Daniel Zhou, Timothee Kheyrkhah, Jeff Smith, and Jonathan Fiscus. Mfc datasets: Large-scale benchmark datasets for media forensic challenge evaluation. In IEEE Winter Applications of Computer Vision Workshops, pages 63\u201372, Jan 2019. 3 [33] David G\u00a8uera and Edward J. Delp. Deepfake video detection", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 116 }, { "text": "using recurrent neural networks. In IEEE International Con- ference on Advanced Video and Signal Based Surveillance, 2018. 3 [34] Rui Huang, Shu Zhang, Tianyu Li, and Ran He. Beyond face rotation: Global and local perception GAN for photorealis- tic and identity preserving frontal view synthesis. In IEEE", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 117 }, { "text": "International Conference on Computer Vision, 2017. 3 [35] Minyoung Huh, Andrew Liu, Andrew Owens, and Alexei A. Efros. Fighting fake news: Image splice detection via learned self-consistency. In European Conference on Com- puter Vision, 2018. 3 [36] Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A.", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 118 }, { "text": "Efros. Image-to-image translation with conditional adver- sarial networks. CVPR, 2017. 5, 14 [37] Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. Progressive Growing of GANs for Improved Quality, Stabil- ity, and Variation. In International Conference on Learning Representations, 2018. 3", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 119 }, { "text": "tian Richardt, Michael Zollh\u00a8ofer, and Christian Theobalt. Deep Video Portraits. ACM Transactions on Graphics 2018 (TOG), 2018. 3, 5 [40] Davis E. King. Dlib-ml: A machine learning toolkit. Journal of Machine Learning Research, 10:1755\u20131758, 2009. 12 [41] Pavel Korshunov and Sebastien Marcel. Deepfakes: a new", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 120 }, { "text": "threat to face recognition? assessment and detection. arXiv preprint arXiv:1812.08685, 2018. 3 [42] Pawel Korus and Jiwu Huang. Multi-scale Analysis Strate- gies in PRNU-based Tampering Localization. IEEE Transac- tions on Information Forensics and Security, 12(4):809\u2013824, Apr. 2017. 3 [43] Guillaume Lample, Neil Zeghidour, Nicolas Usunier, An-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 121 }, { "text": "toine Bordes, and Marc\u2019Aurelio Ranzato Ludovic Denoyer. Fader networks: Manipulating images by sliding attributes. CoRR, abs/1706.00409, 2017. 3 [44] Yuezun Li, Ming-Ching Chang, and Siwei Lyu. In Ictu Oculi: Exposing AI Created Fake Videos by Detecting Eye Blinking. In IEEE WIFS, 2018. 3 [45] Chengjiang Long, Eric Smith, Arslan Basharat, and Anthony", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 122 }, { "text": "Hoogs. A C3D-based Convolutional Neural Network for Frame Dropping Detection in a Single Video Shot. In IEEE Computer Vision and Pattern Recognition Workshops, pages 1898\u20131906, 2017. 3 [46] Yongyi Lu, Yu-Wing Tai, and Chi-Keung Tang. Conditional cyclegan for attribute guided face image generation. In Eu-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 123 }, { "text": "ropean Conference on Computer Vision, 2018. 3 [47] Zhihe Lu, Zhihang Li, Jie Cao, Ran He, and Zhenan Sun. Recent progress of face image synthesis. In IAPR Asian Con- ference on Pattern Recognition, 2017. 3 [48] Patrick Mullan, Davide Cozzolino, Luisa Verdoliva, and Christian Riess. Residual-based forensic comparison of", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 124 }, { "text": "video sequences. In IEEE International Conference on Im- age Processing, 2017. 3 [49] Patrick P\u00b4erez, Michel Gangnet, and Andrew Blake. Pois- son image editing. ACM Transactions on graphics (TOG), 22(3):313\u2013318, 2003. 4, 14 [50] Ramachandra Raghavendra, Kiran B. Raja, Sushma Venkatesh, and Christoph Busch. Transferable Deep-CNN", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 125 }, { "text": "features for detecting digital and print-scanned morphed face images. In IEEE Computer Vision and Pattern Recognition Workshops, 2017. 3 [51] Nicolas Rahmouni, Vincent Nozick, Junichi Yamagishi, and Isao Echizen. Distinguishing computer graphics from nat- ural images using convolution neural networks.", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 126 }, { "text": "In IEEE Workshop on Information Forensics and Security, pages 1\u20136, 2017. 3, 6, 7, 13, 14 [52] Andreas R\u00a8ossler, Davide Cozzolino, Luisa Verdoliva, Chris- tian Riess, Justus Thies, and Matthias Nie\u00dfner. FaceForen- sics: A large-scale video dataset for forgery detection in hu- man faces. arXiv, 2018. 3", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 127 }, { "text": "using an ef\ufb01cient sub-pixel convolutional neural network. In IEEE Conference on Computer Vision and Pattern Recogni- tion, pages 1874\u20131883, 2016. 14 [55] Supasorn Suwajanakorn, Steven M. Seitz, and Ira Kemelmacher-Shlizerman. Synthesizing Obama: learning lip sync from audio. ACM Transactions on Graphics (TOG),", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 128 }, { "text": "expression transfer for facial reenactment. ACM Transac- tions on Graphics (TOG) - Proceedings of ACM SIGGRAPH Asia 2015, 34(6):Art. No. 183, 2015. 2 [59] Justus Thies, Michael Zollh\u00a8ofer, Marc Stamminger, Chris- tian Theobalt, and Matthias Nie\u00dfner. Face2Face: Real-Time Face Capture and Reenactment of RGB Videos. In IEEE", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 129 }, { "text": "Conference on Computer Vision and Pattern Recognition, pages 2387\u20132395, June 2016. 1, 2, 3, 4, 5, 12 [60] Justus Thies, Michael Zollh\u00a8ofer, Marc Stamminger, Chris- tian Theobalt, and Matthias Nie\u00dfner. FaceVR: Real-Time Gaze-Aware Facial Reenactment in Virtual Reality. ACM Transactions on Graphics (TOG), 2018. 3", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 130 }, { "text": "feature interpolation for image content changes. In IEEE Conference on Computer Vision and Pattern Recognition, 2017. 3 [63] Weihong Wang and Hany Farid. Exposing Digital Forgeries in Interlaced and Deinterlaced Video. IEEE Transactions on Information Forensics and Security, 2(3):438\u2013449, 2007. 3 [64] Markos Zampoglou, Symeon Papadopoulos, , and Yiannis", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 131 }, { "text": "Kompatsiaris. Detecting image splicing in the wild (Web). In IEEE International Conference on Multimedia & Expo Workshops (ICMEW), 2015. 3 [65] Kaipeng Zhang, Zhanpeng Zhang, Zhifeng Li, and Yu Qiao. Joint face detection and alignment using multitask cascaded convolutional networks. IEEE Signal Processing Letters,", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 132 }, { "text": "23(10):1499\u20131503, Oct 2016. 14 [66] Peng Zhou, Xintong Han, Vlad I. Morariu, and Larry S. Davis. Two-stream neural networks for tampered face de- tection. In IEEE Computer Vision and Pattern Recognition Workshops, pages 1831\u20131839, 2017. 3 [67] Peng Zhou, Xintong Han, Vlad I. Morariu, and Larry S. Davis. Learning rich features for image manipulation de-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 133 }, { "text": "tection. In CVPR, 2018. 3 [68] Michael Zollh\u00a8ofer, Justus Thies, Darek Bradley, Pablo Garrido, Thabo Beeler, Patrick P\u00b4eerez, Marc Stamminger, Matthias Nie\u00dfner, and Christian Theobalt. State of the art on monocular 3d face reconstruction, tracking, and applica- tions. Computer Graphics Forum, 37(2):523\u2013550, 2018. 1,", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 134 }, { "text": "Validation Test Pristine 366,847 68,511 73,770 DeepFakes 366,835 68,506 73,768 Face2Face 366,843 68,511 73,770 FaceSwap 291,434 54,618 59,640 NeuralTextures 291,834 54,630 59,672 Table 3: Number of images per manipulation method. DeepFakes manipulates every frame of the target sequence, whereas FaceSwap and NeuralTextures only manipulate the", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 135 }, { "text": "minimum number of frames across the source and target video. Face2Face, however, maps all source expressions to the target sequence and rewinds the target video if neces- sary. Number of manipulated frames can vary due to miss- detection in the respective face tracking modules of our ma- nipulation methods.", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 136 }, { "text": "Appendix In FaceForensics++, we evaluate the performance of state-of-the-art facial manipulation detection approaches using a large-scale dataset that we generated with four dif- ferent facial manipulation methods. In addition, we pro- posed an automated benchmark to compare future detec- tion approaches as well as their robustness against unknown", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 137 }, { "text": "post-processing operations such as compression. This supplemental document reports details on our pris- tine data acquisition (Appendix A), ensuring suited input sequences. Appendix B lists the exact numbers of our bi- nary classi\ufb01cation experiments presented in the main paper. Besides binary classi\ufb01cation, the database is also interesting", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 138 }, { "text": "for evaluating manipulation classi\ufb01cation (Appendix C). In Appendix D, we list all chosen hyperparameters of both the manipulation methods as well as the detection techniques. A. Pristine Data Acquisition For a realistic scenario, we chose to collect videos in the wild, more speci\ufb01cally from YouTube. Early experi-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 139 }, { "text": "ments with all manipulation methods showed that the pris- tine videos have to ful\ufb01ll certain criteria. The target face has to be nearly front-facing and without occlusions, to pre- vent the methods from failing or producing strong artifacts (see Fig. 9). We use the YouTube-8m dataset [4] to col- lect videos with the tags \u201cface\u201d, \u201cnewscaster\u201d or \u201cnewspro-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 140 }, { "text": "gram\u201d and also included videos which we obtained from the YouTube search interface with the same tags and ad- ditional tags like \u201cinterview\u201d, \u201cblog\u201d, or \u201cvideo blog\u201d. To ensure adequate video quality, we only downloaded videos that offer a resolution of 480p or higher. For every video, we save its metadata to sort them by properties later on. In", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 141 }, { "text": "order to match the above requirements, we \ufb01rst process all downloaded videos with the Dlib face detector [40], which is based on Histograms of Oriented Gradients (HOG). Dur- ing this step, we track the largest detected face by ensur- ing that the centers of two detections of consecutive frames are pixel-wise close. The histogram-based face tracker was", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 142 }, { "text": "chosen to ensure that the resulting video sequences con- tain little occlusions and, thus, contain easy-to-manipulate faces. Except FaceSwap, all methods need a suf\ufb01ciently large set of image in a target sequence to train on. We select sequences with at least 280 frames. To ensure a high quality video selection and to avoid videos with face occlusions, we", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 143 }, { "text": "perform a manual screening of the clips which resulted in 1,000 video sequences containing 509, 914 images. All examined manipulation methods need a source and a target video. In case of facial reenactment, the expres- sions of the source video are transferred to the target video while retaining the identity of the target person. In contrast,", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 144 }, { "text": "face swapping methods replace the face in the target video with the face in the source video. To ensure high quality face swapping, we select video pairs with similar large faces (considering the bounding box sizes detected by DLib), the same gender of the persons and similar video frame rates. Table 3 lists the \ufb01nal numbers of our dataset for all ma-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 145 }, { "text": "nipulation methods and the pristine data. B. Forgery Detection In this section, we list all numbers from the graphs of the main paper. Table 4 shows the accuracies of the manipulation-speci\ufb01c forgery detectors (i.e., the detectors are trained on the respective manipulation method). In con- trast, Table 5 shows the accuracies of the forgery detectors", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 146 }, { "text": "Raw Compressed 23 Compressed 40 DF F2F FS NT DF F2F FS NT DF F2F FS NT Steg. Features + SVM [27] 99.03 99.13 98.27 99.88 77.12 74.68 79.51 76.94 65.58 57.55 60.58 60.69 Cozzolino et al. [17] 98.83 98.56 98.89 99.88 81.78 85.32 85.69 80.60 68.26 59.38 62.08 62.42 Bayar and Stamm [10] 99.28 98.79 98.98", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 147 }, { "text": "98.78 90.18 94.93 93.14 86.04 80.95 77.30 76.83 72.38 Rahmouni et al. [51] 98.03 98.96 98.94 96.06 82.16 93.48 92.51 75.18 73.25 62.33 67.08 62.59 MesoNet [5] 98.41 97.96 96.07 97.05 95.26 95.84 93.43 85.96 89.52 84.44 83.56 75.74 XceptionNet [14] 99.59 99.61 99.14 99.36 98.85 98.36 98.23 94.5 94.28", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 148 }, { "text": "91.56 93.7 82.11 Table 4: Accuracy of manipulation-speci\ufb01c forgery detectors. We show the results for raw and the compressed datasets of all four manipulation methods (DF: DeepFakes, F2F: Face2Face, FS: FaceSwap and NT: NeuralTextures). Raw Compressed 23 Compressed 40 DF F2F FS NT P DF F2F FS NT P DF", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 149 }, { "text": "F2F FS NT P Steg. Features + SVM [27] 97.96 98.40 91.35 98.56 98.70 68.80 67.69 70.12 69.21 72.98 67.07 48.55 48.68 55.84 56.94 Cozzolino et al. [17] 97.24 98.51 95.93 98.74 99.53 75.51 86.34 76.81 75.34 78.41 75.63 56.01 50.67 62.15 56.27 Bayar and Stamm [10] 99.25 99.04 96.80 99.11 98.92 90.25 93.96", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 150 }, { "text": "87.74 83.69 77.02 86.93 83.66 74.28 74.36 53.87 Rahmouni et al. [51] 94.83 98.25 97.59 96.21 97.34 79.66 87.87 84.34 62.65 79.52 80.36 62.04 59.90 59.99 56.79 MesoNet [5] 99.24 98.35 98.15 97.96 92.04 89.55 88.60 81.24 76.62 82.19 80.43 69.06 59.16 44.81 77.58 XceptionNet [14] 99.29 99.23 98.39 98.64", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 151 }, { "text": "99.64 97.49 97.69 96.79 92.19 95.41 93.36 88.09 87.42 78.06 75.27 Full Image Xception [14] 87.73 83.22 79.29 79.97 81.46 88.00 84.98 82.23 79.60 65.85 84.06 77.56 76.12 66.03 65.09 Table 5: Detection accuracies when trained on all manipulation methods at once and evaluated on speci\ufb01c manipulation methods or pristine data (DF: DeepFakes, F2F: Face2Face, FS: FaceSwap, NT: NeuralTextures, and P: Pristine). The average", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 152 }, { "text": "accuricies are listed in the main paper. Raw Compressed 23 Compressed 40 DF F2F FS NT All DF F2F FS NT All DF F2F FS NT All 10 videos 89.18 76.6 90.89 93.53 92.81 76.06 59.84 81.15 76.73 67.71 64.55 53.99 60.04 65.14 60.55 50 videos 99.52 98.84 97.56 96.67 95.89 92.48 91.33 92.63 85.98 82.89 75.53 66.44", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 153 }, { "text": "74.25 71.48 65.76 100 videos 99.51 99.09 98.64 98.23 97.54 95.39 95.8 95.56 90.09 87.19 83.68 72.69 79.56 73.72 66.81 300 videos 99.59 99.53 98.78 98.73 98.88 97.30 97.41 97.51 92.4 92.65 91.57 86.38 88.35 79.65 76.01 Table 6: Analysis of the training corpus size. Numbers re\ufb02ect the accuracies of the XceptionNet detector trained on single and", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 154 }, { "text": "all manipulation methods (DF: DeepFakes, F2F: Face2Face, FS: FaceSwap, NT: NeuralTextures and All: all manipulation methods). Raw Compressed 23 Compressed 40 DF F2F FS NT P DF F2F FS NT P DF F2F FS NT P Average 77.60 49.60 76.12 32.28 78.19 78.17 50.19 74.80 30.75 75.41 73.18 43.86 64.26 39.07 62.06", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 155 }, { "text": "Desktop PC 80.41 53.73 75.10 34.36 80.12 81.17 51.57 79.51 29.50 78.10 71.71 44.09 62.99 35.71 64.32 Mobile Phone 74.80 44.96 77.11 30.58 76.40 75.47 48.95 70.08 31.97 72.84 74.62 43.62 65.50 41.85 60.00 Table 7: User study result w.r.t. the used device to watch the images (DF: DeepFakes, F2F: Face2Face, FS: FaceSwap, NT:", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 156 }, { "text": "C. Classi\ufb01cation of Manipulation Method To train the XceptionNet classi\ufb01cation network to dis- tinguish between all four manipulation methods and the pristine images, we adapted the \ufb01nal output layer to return \ufb01ve class probabilities. The network is trained on the full dataset containing all pristine and manipulated images. On", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 157 }, { "text": "raw data the network is able to achieve a 99.03% accuracy, which slightly decreases for the high quality compression to 95.42% and to 80.49% on low quality images. D. Hyperparameters For reproducibility, we detail the hyperparameters used for the methods in the main paper. We structured this sec- tion into two parts, one for the manipulation methods and", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 158 }, { "text": "the second part for the classi\ufb01cation approaches used for forgery detection. D.1. Manipulation Methods DeepFakes and NeuralTextures are learning-based, for the other manipulation methods we used the default param- eters of the approaches. DeepFakes: Our DeepFakes implementation is based on the deepfakes faceswap github project [1]. MTCNN ([65])", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 159 }, { "text": "is used to extract and align the images for each video. Speci\ufb01cally, the largest face in the \ufb01rst frame of a sequence is detected and tracked throughout the whole video. This tracking information is used to extract the training data for DeepFakes. The auto-encoder takes input images of 64 (de- fault). It uses a shared encoder consisting of four convolu-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 160 }, { "text": "tional layers which downsizes the image to a bottleneck of 4 \u00d7 4, where we \ufb02atten the input, apply a fully connected layer, reshape the dense layer and apply a single upscaling using a convolutional layer as well as a pixel shuf\ufb02e layer (see [54]). The two decoders use three identical up-scaling layers to attain full input image resolution. All layers use", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 161 }, { "text": "Leaky ReLus as non-linearities. The network is trained using Adam with a learning rate of 10\u22125, \u03b21 = 0.5 and \u03b22 = 0.999 as well as a batch size of 64. In our experi- ments, we run the training for 200000 iterations on a cloud platform. By exchanging the decoder of one person to an- other, we can generate an identity-swapped face region. To", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 162 }, { "text": "insert the face into the target image, we chose Poisson Im- age Editing [49] to achieve a seamless blending result. NeuralTextures: NeuralTextures is based on a U-Net ar- chitecture. For data generation, we employ the original pipeline and network architecture (for details see [57]). In addition to the photo-metric consistency, we added an ad-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 163 }, { "text": "versarial loss. This adversarial loss is based on the patch- based discriminator used in Pix2Pix [36]. During training we weight the photo-metric loss with 1 and the adversarial loss with 0.001. We train three models per manipulation for a \ufb01xed 45 epochs using the Adam optimizer (with default settings) and manually choose the best performing model", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 164 }, { "text": "based on visual quality. All manipulations are created at a resolution of 512 \u00d7 512 as in the original paper, with a texture resolution of 512 \u00d7 512 and 16 feature per texel. In- stead of using the entire image, we only train and modify the cropped image containing the face bounding box ensur- ing high resolution outputs even on higher resolution im-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 165 }, { "text": "ages. To do so, we enlarge the bounding box obtained by the Face2Face tracker by a factor of 1.8. D.2. Classi\ufb01cation Methods For our forgery detection pipeline proposed in the main paper, we conducted studies with \ufb01ve classi\ufb01cation ap- proaches based on convolutional neural networks. The net- works are trained using the Adam optimizer with different", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 166 }, { "text": "parameters for learning-rate and batch-size. In particular, for the network proposed in Cozzolino at al. [17] the used learning-rate is 10\u22125 with batch-size 16. For the proposal of Bayar and Stamm [10], we use a learning-rate equal to 10\u22125 with a batch-size of 64. The network proposed by Rahmouni [51] is trained with a learning-rate of 10\u22124 and a", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 167 }, { "text": "batch-size equal to 64. MesoNet [5] uses a batch-size of 76 and the learning-rate is set to 10\u22123. Our XceptionNet [14]- based approach is trained with a learning-rate of 0.0002 and a batch-size of 32. All detection methods are trained with the Adam optimizer using the default values for the mo- ments (\u03b21 = 0.9, \u03b22 = 0.999, \u03f5 = 10\u22128).", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 168 }, { "text": "We compute validation accuracies ten times per epoch and stop the training process if the validation accuracy does not change for 10 consecutive checks. Validation and test accuracies are computed on 100 images per video, training is evaluated on 270 images per video to account for frame count imbalance in our videos. Finally, we solve the imbal-", "source": "FaceForensics++", "year": 2019, "url": "https://arxiv.org/abs/1901.08971", "id": 169 }, { "text": "1 DeepFakes and Beyond: A Survey of Face Manipulation and Fake Detection Ruben Tolosana, Ruben Vera-Rodriguez, Julian Fierrez, Aythami Morales and Javier Ortega-Garcia Biometrics and Data Pattern Analytics - BiDA Lab, Universidad Autonoma de Madrid, Spain {ruben.tolosana, ruben.vera, julian.\ufb01errez, aythami.morales, javier.ortega}@uam.es", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 170 }, { "text": "Abstract\u2014The free access to large-scale public databases, together with the fast progress of deep learning techniques, in particular Generative Adversarial Networks, have led to the generation of very realistic fake content with its corresponding implications towards society in this era of fake news.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 171 }, { "text": "This survey provides a thorough review of techniques for manipulating face images including DeepFake methods, and methods to detect such manipulations. In particular, four types of facial manipulation are reviewed: i) entire face synthesis, ii) identity swap (DeepFakes), iii) attribute manipulation, and iv)", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 172 }, { "text": "expression swap. For each manipulation group, we provide details regarding manipulation techniques, existing public databases, and key benchmarks for technology evaluation of fake detection methods, including a summary of results from those evaluations. Among all the aspects discussed in the survey, we pay special", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 173 }, { "text": "attention to the latest generation of DeepFakes, highlighting its improvements and challenges for fake detection. In addition to the survey information, we also discuss open issues and future trends that should be considered to advance in the \ufb01eld. Index Terms\u2014Fake News, DeepFakes, Media Forensics, Face", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 174 }, { "text": "Manipulation, Face Recognition, Biometrics, Databases, Bench- mark I. INTRODUCTION F AKE images and videos including facial information generated by digital manipulation, in particular with DeepFake methods [1], have become a great public concern recently [2], [3]. The very popular term \u201cDeepFake\u201d is referred", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 175 }, { "text": "to a deep learning based technique able to create fake videos by swapping the face of a person by the face of another person. This term was originated after a Reddit user named \u201cdeepfakes\u201d claimed in late 2017 to have developed a machine learning algorithm that helped him to transpose celebrity faces", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 176 }, { "text": "into porn videos [4]. In addition to fake pornography, some of the more harmful usages of such fake content include fake news, hoaxes, and \ufb01nancial fraud. As a result, the area of research traditionally dedicated to general media forensics [5]\u2013 [11], is being invigorated and is now dedicating growing ef-", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 177 }, { "text": "forts for detecting facial manipulation in image and video [12]. Part of these renewed efforts in fake face detection are built around past research in biometric anti-spoo\ufb01ng [13]\u2013[15] and modern data-driven deep learning [16], [17]. The growing interest in fake face detection is demonstrated through the", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 178 }, { "text": "increasing number of workshops in top conferences [18]\u2013 [22], international projects such as MediFor funded by the Defense Advanced Research Project Agency (DARPA), and competitions such as the recent Media Forensics Challenge (MFC2018)1 and the Deepfake Detection Challenge (DFDC)2 launched by the National Institute of Standards and Technol-", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 179 }, { "text": "ogy (NIST) and Facebook, respectively. Traditionally, the number and realism of facial manipula- tions have been limited by the lack of sophisticated editing tools, the domain expertise required, and the complex and time-consuming process involved. For example, an early work in this topic [23] was able to modify the lip motion of a person", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 180 }, { "text": "speaking using a different audio track, by making connections between the sounds of the audio track and the shape of the subject\u2019s face. However, from these early works up to date, many things have rapidly evolved in the last years. Nowadays, it is becoming increasingly easy to automatically synthesise", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 181 }, { "text": "non-existent faces or manipulate a real face of one person in an image/video, thanks to: i) the accessibility to large-scale public data, and ii) the evolution of deep learning techniques that eliminate many manual editing steps such as Autoencoders (AE) and Generative Adversarial Networks (GAN) [24], [25].", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 182 }, { "text": "As a result, open software and mobile application such as ZAO3 and FaceApp4 have been released opening the door to anyone to create fake images and videos, without any experience in the \ufb01eld needed. In response to those increasingly sophisticated and realistic manipulated content, large efforts are being carried out by", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 183 }, { "text": "the research community to design improved methods for face manipulation detection. Traditional fake detection methods in media forensics have been commonly based on: i) in-camera \ufb01ngerprints, the analysis of the intrinsic \ufb01ngerprints introduced by the camera device, both hardware and software, such as", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 184 }, { "text": "the optical lens [27], colour \ufb01lter array and interpolation [28], [29], and compression [30], [31], among others, and ii) out- camera \ufb01ngerprints, the analysis of the external \ufb01ngerprints introduced by editing software, such as copy-paste or copy- move different elements of the image [32], [33], reduce the", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 185 }, { "text": "frame rate in a video [34]\u2013[36], etc. However, most of the features considered in traditional fake detection methods are highly dependent on the speci\ufb01c training scenario, being therefore not robust against unseen conditions [6], [8], [16]. This is of special importance in the era we live in as most", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 186 }, { "text": "media fake content is usually shared on social networks, whose platforms automatically modify the original image/video, for example, through compression and resize operations [12]. 1https://www.nist.gov/itl/iad/mig/media-forensics-challenge-2018 2https://deepfakedetectionchallenge.ai/ 3https://apps.apple.com/cn/app/id1465199127", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 187 }, { "text": "2 Real Entire Face Synthesis Fake Identity Swap Expression Swap Attribute Manipulation Fake Real Fake Real Source Target Fake Real Source Target Fig. 1. Real and fake examples of each facial manipulation group. For Entire Face Synthesis, real images are extracted from http://www.whichfaceisreal.com/ and", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 188 }, { "text": "fake images from https://thispersondoesnotexist.com. For Identity Swap, face images are extracted from Celeb-DF database [26]. For Attribute Manipulation, real images are extracted from http://www.whichfaceisreal.com/ and fake images are generated using FaceApp. Finally, for Expression Swap, images are", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 189 }, { "text": "extracted from FaceForensics++ [12]. This survey provides an in-depth review of digital manip- ulation techniques applied to facial content due to the large number of possible harmful applications, e.g., the generation of fake news that would provide misinformation in political elections and security threats [37], [38]. Speci\ufb01cally, we cover", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 190 }, { "text": "four types of manipulations: i) entire face synthesis, ii) identity swap, iii) attribute manipulation, and iv) expression swap. These four main types of face manipulation are well estab- lished by the research community, receiving most attention in the last few years. Besides, we also review in this survey some", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 191 }, { "text": "other challenging and dangerous face manipulation techniques that are not so popular yet like face morphing. Finally, for completeness, we would like to highlight other recent surveys in the \ufb01eld. In [39], the authors cover the topic of DeepFakes from a general perspective, proposing the R.E.A.L framework to manage DeepFake risks. In addition,", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 192 }, { "text": "Verdoliva has recently surveyed in [40] traditional manipu- lation and fake detection approaches considered in general media forensics, and also the latest deep learning techniques. The present survey complements [39] and [40] with a more detailed review of each facial manipulation group, including", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 193 }, { "text": "manipulation techniques, existing public databases, and key benchmarks for technology evaluation of fake detection meth- ods, including a summary of results from those evaluations. In addition, we pay special attention to the latest generation of DeepFakes, highlighting its improvements and challenges", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 194 }, { "text": "for fake detection. The remainder of the article is organised as follows. We \ufb01rst provide in Sec. II a general description of different types of facial manipulation. Then, from Sec. III to Sec. VI we describe the key aspects of each type of facial manipulation including public databases for research, detection methods,", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 195 }, { "text": "3 II. TYPES OF FACIAL MANIPULATIONS Facial manipulations can be categorised in four main dif- ferent groups regarding the level of manipulation. Fig. 1 graphically summarises each facial manipulation group. A description of each of them is provided below, from higher to lower level of manipulation:", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 196 }, { "text": "\u2022 Entire Face Synthesis: this manipulation creates entire non-existent face images, usually through powerful GAN, e.g., through the recent StyleGAN approach proposed in [41]. These techniques achieve astonishing results, generating high-quality facial images with a high level of realism. Fig. 1 shows some examples for entire face", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 197 }, { "text": "synthesis generated using StyleGAN5. This manipulation could bene\ufb01t many different sectors such as the video game and 3D-modelling industries, but it could also be used for harmful applications such as the creation of very realistic fake pro\ufb01les in social networks in order to generate misinformation.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 198 }, { "text": "\u2022 Identity Swap: this manipulation consists of replacing the face of one person in a video with the face of another person. Two different approaches are usually considered: i) classical computer graphics-based techniques such as FaceSwap6, and ii) novel deep learning techniques known as DeepFakes7, e.g., the recent ZAO mobile application.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 199 }, { "text": "Very realistic videos of this type of manipulation can be seen on Youtube8. This type of manipulation could bene\ufb01t many different sectors, in particular the \ufb01lm industry. However, in the other side, it could also be used for bad purposes such as the creation of celebrity pornographic videos, hoaxes, and \ufb01nancial fraud, among many others.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 200 }, { "text": "\u2022 Attribute Manipulation: this manipulation, also known as face editing or face retouching, consists of modifying some attributes of the face such as the colour of the hair or the skin, the gender, the age, adding glasses, etc [42]. This manipulation process is usually carried out through GAN such as the StarGAN approach proposed in [43].", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 201 }, { "text": "One example of this type of manipulation is the popular FaceApp mobile application. Consumers could use this technology to try on a broad range of products such as cosmetics and makeup, glasses, or hairstyles in a virtual environment. \u2022 Expression Swap: this manipulation, also known as face reenactment, consists of modifying the facial expression", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 202 }, { "text": "of the person. Although different manipulation techniques are proposed in the literature, e.g., at image level through popular GAN architectures [44], in this group we focus on the most popular techniques Face2Face and Neural- Textures [45], [46], which replaces the facial expression of one person in a video with the facial expression of", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 203 }, { "text": "another person. This type of manipulation could be used with serious consequences, e.g., the popular video of Mark Zuckerberg saying things he never said9. 5https://thispersondoesnotexist.com 6https://github.com/MarekKowalski/FaceSwap 7https://github.com/deepfakes/faceswap 8https://www.youtube.com/watch?v=UlvoEW7l5rs", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 204 }, { "text": "9https://www.bbc.com/news/technology-48607673 TABLE I ENTIRE FACE SYNTHESIS: PUBLICLY AVAILABLE DATABASES. Database Real Images Fake Images 100K-Generated-Images (2019) [41] - 100,000 (StyleGAN) 100K-Faces (2019) [47] - 100,000 (StyleGAN) DFFD (2020) [17] - 100,000 (StyleGAN) 200,000 (ProGAN) iFakeFaceDB (2020)", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 205 }, { "text": "of relevance here, all of them based on the same GAN architectures: ProGAN [48] and StyleGAN [41]. It is inter- esting to remark that each fake image may be characterised by a speci\ufb01c GAN \ufb01ngerprint just like natural images are identi\ufb01ed by a device-based \ufb01ngerprint (i.e., PRNU). In fact, these \ufb01ngerprints seem to be dependent not only of the GAN", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 206 }, { "text": "architecture, but also of the different instances of it [49]\u2013[51]. In addition, as indicated in Table I, it is important to note that the four mentioned databases only contain fake images generated using the GAN architectures discussed. In order to perform fake detection experiments on this manipulation", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 207 }, { "text": "group, researchers need to obtain real face images from other public databases such as CelebA [52], FFHQ [41], CASIA- WebFace [53], and VGGFace2 [54], among others. We provide next a description of each public database. In [41], Karras et al. released a set of 100,000 synthetic face images, named 100K-Generated-Images10. This database", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 208 }, { "text": "was generated using their proposed StyleGAN architecture, which was trained using the FFHQ dataset [41]. StyleGAN is an improved version of their previous popular approach ProGAN, which introduced a new training methodology based on improving both generator and discriminator progressively. StyleGAN proposes an alternative generator architecture that", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 209 }, { "text": "leads to an automatically learned, unsupervised separation of high-level attributes (e.g., pose and identity when trained on human faces) and stochastic variation in the generated images (e.g., freckles, hair), and it enables intuitive, scale-speci\ufb01c control of the synthesis. Another public database is 100K-Faces [47]. This database", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 210 }, { "text": "contains 100,000 synthetic images generated using Style- GAN. In this database, contrary to the 100K-Generated- Images database, the StyleGAN network was trained using around 29,000 photos from 69 different models, considering face images from a more controlled scenario (e.g., with a \ufb02at background). Thus, no strange artifacts created by the", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 211 }, { "text": "4 (a) Fake (b) Fake after GANprintR Fig. 2. Examples of a fake image created using StyleGAN and its improved version after removing the GAN-\ufb01ngerprint information with GANprintR [16]. Recently, Dang et al. introduced in [17] a new database named Diverse Fake Face Dataset (DFFD). Regarding the entire face synthesis manipulation, the authors created 100,000", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 212 }, { "text": "and 200,000 fake images through the pre-trained ProGAN and StyleGAN models, respectively. Finally, Neves et al. presented in [16] the iFakeFaceDB database. This database comprises 250,000 and 80,000 syn- thetic face images created with StyleGAN and ProGAN, respectively. As an additional feature in comparison to pre-", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 213 }, { "text": "vious databases, and in order to hinder fake detectors, in this database the \ufb01ngerprints produced by the GAN architectures were removed through an approach named GANprintR (GAN \ufb01ngerprint Removal), while keeping very realistic appearance. Fig. 2 shows an example of a fake image directly generated with StyleGAN and its improved version after removing the", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 214 }, { "text": "GAN-\ufb01ngerprint information. As a result of the GANprintR step, iFakeFaceDB presents a higher challenge for advanced fake detectors compared with the other databases. B. Manipulation Detection Different studies have recently evaluated the dif\ufb01culty of detecting whether faces are real of arti\ufb01cially generated.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 215 }, { "text": "cases, different evaluation metrics are considered, e.g., Area Under the Curve (AUC) or Equal Error Rate (EER), which complicates the comparison among the studies. Some authors propose to analyse the internal GAN pipeline in order to detect different artifacts between real and fake images. In [55], the authors hypothesised that the colour", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 216 }, { "text": "is markedly different between real camera images and fake synthesis images. They proposed a detection system based on colour features and a linear Support Vector Machine (SVM) for the \ufb01nal classi\ufb01cation, achieving a \ufb01nal 70.0% AUC for the best performance when evaluating with the NIST MFC2018 dataset [62].", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 217 }, { "text": "Another interesting approach in this line was proposed in [56]. Wang et al. conjectured that monitoring neuron behavior could also serve as an asset in detecting fake faces since layer-by-layer neuron activation patterns may capture more subtle features that are important for the facial ma- nipulation detection system. Their proposed approach, named", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 218 }, { "text": "FakeSpoter, extracted as features neuron coverage behaviors of real and fake faces from deep face recognition systems (i.e., VGG-Face [63], OpenFace [64], and FaceNet [65]), and then trained a SVM for the \ufb01nal classi\ufb01cation. The authors tested their proposed approach using real faces from CelebA-HQ [48]", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 219 }, { "text": "and FFHQ [41] databases and synthetic faces created through InterFaceGAN [66] and StyleGAN [41], achieving for the best performance a \ufb01nal 84.7% fake detection accuracy using the FaceNet model. Better results have been recently reported in [57]. The authors proposed a fake detection system based on the analysis", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 220 }, { "text": "of the convolutional traces. Features were extracted using the Expectation Maximization algorithm [67]. Popular classi\ufb01ers such as k-Nearest Neighbours (k-NN), SVM, and Linear Discriminant Analysis (LDA) were used for the \ufb01nal detection. Their proposed approach was tested using fake images gen- erated through AttGAN [68], GDWCT [69], StarGAN [43],", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 221 }, { "text": "StyleGAN, and StyleGAN2 [70], achieving a \ufb01nal 99.81% Acc. for the best performance. Fake detection systems inspired in steganalysis have also been studied. Nataraj et al. proposed in [58] a detection system based on a combination of pixel co-occurrence ma- trices and Convolutional Neural Networks (CNN). Their pro-", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 222 }, { "text": "posed approach was initially tested through a database of various objects and scenes created through CycleGAN [71]. Besides, the authors performed an interesting analysis to see the robustness of the proposed approach against fake images created through different GAN architectures (CycleGAN vs. StarGAN), with good generalisation results. This detection", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 223 }, { "text": "approach was implemented later on in [16] considering images from the 100K-Faces database, achieving an EER of 12.3% for the best fake detection performance. This result is remarked in italics in Table II to indicate that it was not provided in the original paper. Many studies have also focused on the detection of the spe-", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 224 }, { "text": "cial \ufb01ngerprints inserted by GAN architectures using pure deep learning methods. Yu et al. proposed in [59] an attribution net- work architecture to map an input image to its corresponding \ufb01ngerprint image. Therefore, they learned a model \ufb01ngerprint for each source (each GAN instance plus the real world), such", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 225 }, { "text": "that the correlation index between one image \ufb01ngerprint and each model \ufb01ngerprint serves as softmax logit for classi\ufb01ca- tion. Their proposed approach was tested using real faces from CelebA database [52] and synthetic faces created through dif- ferent GAN approaches (ProGAN [48], SNGAN [72], Cramer-", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 226 }, { "text": "GAN [73], and MMDGAN [74]), achieving a \ufb01nal 99.5% fake detection accuracy for the best performance. However, this approach seemed not to be very robust against unseen simple image perturbation attacks such as noise, blur, cropping or compression, unless the models were re-trained again. Related to the unseen conditions just commented, Marra et", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 227 }, { "text": "al. performed in [60] an interesting study in order to detect unseen types of fake generated data. Concretely, they proposed a multi-task incremental learning detection method in order to detect and classify new types of GAN generated images, without worsening the performance on the previous ones. Two different solutions regarding the position of the classi\ufb01er", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 228 }, { "text": "5 TABLE II ENTIRE FACE SYNTHESIS: COMPARISON OF DIFFERENT STATE-OF-THE-ART DETECTION APPROACHES. THE BEST RESULTS ACHIEVED FOR EACH PUBLIC DATABASE ARE REMARKED IN BOLD. RESULTS IN italics INDICATE THAT THEY WERE NOT PROVIDED IN THE ORIGINAL WORK. AUC = AREA UNDER THE CURVE, ACC. = ACCURACY, EER = EQUAL ERROR RATE.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 229 }, { "text": "Study Method Classi\ufb01ers Best Performance Databases (Generation) McCloskey and Albright (2018) [55] GAN-Pipeline Features SVM AUC = 70.0% NIST MFC2018 Wang et al. (2019) [56] GAN-Pipeline Features SVM Acc. = 84.7% Own (InterFaceGAN, StyleGAN) Guarnera et al. (2020) [57] GAN-Pipeline Features k-NN, SVM, LDA", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 230 }, { "text": "Acc. = 99.81% Own (AttGAN, GDWCT, StarGAN, StyleGAN, StyleGAN2) Nataraj et al. (2019) [58] Steganalysis Features CNN EER = 12.3% [16] 100K-Faces (StyleGAN) Yu et al. (2019) [59] Deep Learning Features CNN Acc. = 99.5% Own (ProGAN, SNGAN, CramerGAN, MMDGAN) Marra et al. (2019) [60] Deep Learning Features", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 231 }, { "text": "CNN + Incremental Learning Acc. = 99.3% Own (CycleGAN, ProGAN, Glow, StarGAN, StyleGAN) Dang et al. (2020) [17] Deep Learning Features CNN + Attention Mechanism AUC = 100% EER = 0.1% DFFD (ProGAN, StyleGAN) Neves et al. (2020) [16] Deep Learning Features CNN EER = 0.3% 100K-Faces (StyleGAN) EER = 4.5%", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 232 }, { "text": "iFakeFaceDB Hulzebosch et al. (2020) [61] Deep Learning Features CNN, AE Acc. = 99.8% Own (StarGAN, Glow, ProGAN, StyleGAN) incremental learning [75]: i) Multi-Task MultiClassi\ufb01er (MT- MC), and ii) Multi-Task Single Classi\ufb01er (MT-SC). Regarding the experimental framework, \ufb01ve different GAN approaches", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 233 }, { "text": "were considered in the study, CycleGAN [71], ProGAN [48], Glow [76], StarGAN [43], and StyleGAN [41]. Their proposed detection approach, based on the XceptionNet model, achieved promising results being able to correctly detect new GAN generated images. Attention mechanisms have also been applied to further", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 234 }, { "text": "improve the training process of the detection systems. Dang et al. carried out in [17] a complete analysis of different types of facial manipulations. They proposed to use attention mechanisms and popular CNN models such as Xception- Net and VGG16. For the entire face synthesis manipulation, the authors achieved a \ufb01nal 100% AUC and around 0.1%", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 235 }, { "text": "EER considering real faces from CelebA [52], FFHQ [41], and FaceForensics++ [12] databases and fake images created through ProGAN [48] and StyleGAN [41] approaches. The impressive results achieved show the importance of novel attention mechanisms [77]. Neves et al. performed in [16] an in-depth experimental", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 236 }, { "text": "assessment of this type of facial manipulation considering different state-of-the-art detection systems and experimental conditions, i.e., controlled and in-the-wild scenarios. Four different fake databases were considered: i) 150,000 fake faces collected online11 and based on StyleGAN architecture, ii)", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 237 }, { "text": "the 100K-faces public database, iii) 80,000 synthetic faces generated using ProGAN, and iv) the iFakeFaceDB database, an improved version of previous fake databases in which the GAN-\ufb01ngerprint information has been removed using the GANprintR approach. In controlled scenarios, they achieved 11https://thispersondoesnotexist.com", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 238 }, { "text": "similar results as the best previous studies (EER = 0.02%). However, in more challenging scenarios in which images (real and fake) come from different sources (mismatch of datasets), a high degradation of the fake detection performance is observed. Finally, the results achieved over their public iFakeFaceDB database with an EER = 4.5% for the best fake", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 239 }, { "text": "detectors remark how challenging is iFakeFaceDB even for the most advanced manipulation detection methods. Related to this enhanced fake content, Cozzolino et al. proposed in [78] a similar approach based on GAN to inject camera traces into synthetic images to spoof state-of-the-art fake detectors.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 240 }, { "text": "Similar to [16], Hulzebosch et al. have recently performed in [61] an in-depth analysis of this face manipulation consid- ering different scenarios such as cross-model, cross-data, and post-processing. Fake detection approaches were based on the popular Xception network and ForensicTransfer [79], which", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 241 }, { "text": "is an Autoencoder approach. In general, bad generalisation results were obtained under unseen scenarios, similar to [16]. Finally, we also include for completeness some important references to other recent studies focused on the detection of general GAN-based image manipulations, not facial ones. In", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 242 }, { "text": "particular, we refer the reader to [80], [81]. IV. IDENTITY SWAP A. Manipulation Techniques and Public Databases This is one of the most popular face manipulation research lines nowadays due to the great public concerns around DeepFakes [2], [3]. It consists of replacing the face of one person in a video with the face of another person. Unlike", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 243 }, { "text": "6 Since publicly available fake databases such as the UADFV database [82], up to the recent Celeb-DF and DFDC databases [26], [83], many visual improvements have been carried out, increasing the realism of fake videos. As a result, identity swap databases can be divided into two different generations. Table III summarises the main details of each", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 244 }, { "text": "public database, grouped in each generation. As can be seen, in this type of facial manipulation both real and fake videos are usually included in the databases. In this section, we \ufb01rst provide the main details of each database, to \ufb01nally summarise at a higher level the key differences among the two generations.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 245 }, { "text": "Three different databases are grouped in the \ufb01rst generation. UADFV was one of the \ufb01rst public databases [82]. This database comprises 49 real videos from Youtube, which were used to create 49 fake videos through the FakeApp mobile application12, swapping in all of them the original face with the face of Nicolas Cage. Therefore, only one identity is", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 246 }, { "text": "considered in all fake videos. Each video represents one individual, with a typical resolution of 294\u00d7500 pixels, and 11.14 seconds on average. Korshunov and Marcel introduced in [1] the Deepfake- TIMIT database. This database comprises 620 fake videos of 32 subjects from the VidTIMIT database [84]. Fake videos", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 247 }, { "text": "were created using the public GAN-based face-swapping al- gorithm13. In that approach, the generative network is adopted from CycleGAN [71], using the weights of FaceNet [65]. The method Multi-Task Cascaded Convolution Networks is used for more stable detections and reliable face alignment [85]. Besides, the Kalman \ufb01lter is also considered to smooth the", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 248 }, { "text": "bounding box positions over frames and eliminate jitter on the swapped face. Regarding the scenarios considered in DeepfakeTIMIT, two different qualities are considered: i) low quality (LQ) with images of 64\u00d764 pixels, and ii) high quality (HQ) with images of 128\u00d7128 pixels. Additionally, different", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 249 }, { "text": "blending techniques were applied to the fake videos regarding the quality level. One of the most popular databases in this type of facial manipulation is FaceForensics++ [12]. This database was introduced early 2019 as an extension of the original Face- Forensics database [87], which was focused only on expression", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 250 }, { "text": "swap. FaceForensics++ contains 1000 real videos extracted from Youtube. Regarding the identity swap fake videos, they were generated using both computer graphics and DeepFake approaches (i.e., learning approach). For the computer graph- ics approach, the authors considered the publicly available FaceSwap algorithm14 whereas for the DeepFake approach,", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 251 }, { "text": "fake videos were created through the DeepFake FaceSwap GitHub implementation15. The FaceSwap approach consists of face alignment, Gauss Newton optimization and image blending to swap the face of the source person to the target person. The DeepFake approach, as indicated in [12], is based on two autoencoders with a shared encoder that are trained", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 252 }, { "text": "to reconstruct training images of the source and the target 12https://www.malavida.com/en/soft/fakeapp/ 13https://github.com/shaoanlu/faceswap-GAN 14https://github.com/MarekKowalski/FaceSwap 15https://github.com/deepfakes/faceswap TABLE III IDENTITY SWAP: PUBLICLY AVAILABLE DATABASES. 1st Generation", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 253 }, { "text": "Database Real Videos Fake Videos UADFV (2018) [82] 49 (Youtube) 49 (FakeApp) DeepfakeTIMIT (2018) [1] - 620 (faceswap-GAN) FaceForensics++ (2019) [12] 1,000 (Youtube) 1,000 (FaceSwap) 1,000 (DeepFake) 2nd Generation Database Real Videos Fake Videos DeepFakeDetection (2019) [86] 363 (Actors) 3,068 (DeepFake)", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 254 }, { "text": "Celeb-DF (2019) [26] 890 (Youtube) 5,639 (DeepFake) DFDC Preview (2019) [83] 1,131 (Actors) 4,119 (Unknown) face, respectively. A face detector is used to crop and to align the images. To create a fake image, the trained encoder and decoder of the source face are applied to the target face. The autoencoder output is then blended with the rest of", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 255 }, { "text": "the image using Poisson image editing [88]. Regarding the \ufb01gures of the FaceForensics++ database, 1000 fake videos were generated for each approach. Later on, a new dataset named DeepFakeDetection, grouped inside the 2nd generation due to its higher realism, was included in the FaceForensics++ framework with the support of Google", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 256 }, { "text": "are considered, in particular: i) RAW (original quality), ii) HQ (constant rate quantization parameter equal to 23), and iii) LQ (constant rate quantization parameter equal to 40). This aspect simulates the video processing techniques usually applied in social networks. Regarding the databases included in the 2nd generation, we", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 257 }, { "text": "highlight the recent Celeb-DF and DFDC databases released at the end of 2019. Li et al. presented in [26] the Celeb- DF database. This database aims to provide fake videos of better visual qualities, similar to the popular videos that are shared on the Internet16, in comparison to previous databases", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 258 }, { "text": "that exhibit low visual quality with many visible artifacts. Celeb-DF consists of 890 real videos extracted from Youtube, and 5,639 fake videos, which were created through a re\ufb01ned version of a public DeepFake generation algorithm, improving aspects such as the low resolution of the synthesised faces and", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 259 }, { "text": "colour inconsistencies. Facebook in collaboration with other companies and aca- demic institutions such as Microsoft, Amazon, and the MIT launched at the end of 2019 a new challenge named the Deep- fake Detection Challenge (DFDC) [83]. They \ufb01rst released a preview dataset consisting of 1,131 real videos from 66 paid", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 260 }, { "text": "7 st Identity Swap: 1 Generation nd Identity Swap: 2 Generation Low-Quality Synthesised Faces Colour Contrast in the Fake Mask Visible Elements from Original Video Strange Artifacts between Frames Visible Boundaries in the Fake Mask High Pose Variations Scenarios: Indoors and Outdoors Light Conditions: Day, Night, etc.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 261 }, { "text": "Distance from the Camera Weaknesses that limit the naturalness and facilitate fake detection Improvements that augment the naturalness and hinder fake detection Fig. 3. Graphical representation of the weaknesses present in identity swap databases of the 1st generation and the improvements carried out in the 2nd", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 262 }, { "text": "8 actors, and 4,119 fake videos. Fake videos were generated using two different unknown approaches. The complete DFDC dataset was released later and comprises over 470 GB of content (real and fake)17. Finally, to conclude this section, we discuss at a higher level the key differences among fake databases from the 1st and 2nd", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 263 }, { "text": "generations. In general, fake videos from the 1st generation are characterised by: i) low-quality synthesised faces, ii) different colour contrast among the synthesised fake mask and the skin of the original face, iii) visible boundaries of the fake mask, iv) visible facial elements from the original video, v) low", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 264 }, { "text": "pose variations, and vi) strange artifacts among sequential frames. Also, they usually consider controlled scenarios in terms of camera position and light conditions. Many of these aspects have been successfully improved in databases of the 2nd generation, not only at visual level, but also in terms", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 265 }, { "text": "of variability (in-the-wild scenarios). For example, the recent DFDC database considers different acquisition scenarios (i.e., indoors and outdoors), light conditions (i.e., day, night, etc.), distances from the person to the camera, and pose variations, among others. Fig. 3 graphically summarises the weaknesses", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 266 }, { "text": "present in identity swap databases of the 1st generation and the improvements carried out in the 2nd generation. Finally, it also interesting to remark the larger number of fake videos included in the databases of the 2nd generation. B. Manipulation Detection The development of novel methods to detect identity swap", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 267 }, { "text": "manipulations is continuously evolving. Table IV provides a comparison of the most relevant detection approaches in this area. For each study we include information related to the method, classi\ufb01ers, best performance, and databases for research. We highlight in bold the best results achieved for each public database. It is important to remark that in some", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 268 }, { "text": "cases, different evaluation metrics are considered (e.g., AUC and EER), which complicates the comparison among studies. Finally, the results highlighted in italics indicate the gener- alisation capacity of the detection systems against different unseen databases, i.e., those databases were not considered", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 269 }, { "text": "for training. These results have been extracted from [26] and were not included in the original publications. The \ufb01rst studies in this area focused on the audio-visual artifacts existed in the 1st generation of fake videos. Korshunov and Marcel evaluated in [1] baseline approaches based on the inconsistencies between lip movements and audio speech,", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 270 }, { "text": "as well as several variations of image-based systems often used in biometrics. For the \ufb01rst case, they considered Mel- Frequency Cepstral Coef\ufb01cients (MFCCs) as audio features and distances between mouth landmarks as visual features. Principal Component Analysis (PCA) was then used to reduce the dimensionality of the blocks of features, and \ufb01nally", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 271 }, { "text": "Recurrent Neural Networks (RNNs) based on Long Short- Term Memory (LSTM) to detect real of fake videos (based on [101]). For the second case, they evaluated detection approaches based on: i) raw faces as features, and ii) image quality measures (IQM) [102]. In particular, they used a set 17https://www.kaggle.com/c/deepfake-detection-challenge", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 272 }, { "text": "of 129 features related to measures like signal to noise ratio, specularity, blurriness, etc. PCA with LDA, or SVM were considered for the \ufb01nal classi\ufb01cation. Their proposed detection approach based on IQM+SVM provided the best results, with a \ufb01nal 3.3% and 8.9% EER for the LQ and HQ scenarios of the DeepfakeTIMIT database, respectively.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 273 }, { "text": "In this line, Matern et al. proposed in [89] fake detection systems based on relatively simple visual aspects such as eye colour, missing re\ufb02ections, and missing details in the eye and teeth areas. Two different classi\ufb01ers were considered in this analysis: i) a logistic regression model, and ii) a Multilayer", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 274 }, { "text": "Perceptron (MLP) [103]. Their proposed approach was tested using a private database, achieving a \ufb01nal 85.1% AUC for the MLP system. Fake detection systems based on facial expressions and head movements have also been proposed in the literature. Yang et al. observes in [90] that some DeepFakes are created by", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 275 }, { "text": "splicing synthesised face regions into the original image, and in doing so, introducing errors that can be revealed when 3D head poses are estimated from the face images. Thus, they performed an study based on the differences between head poses estimated using a full set of facial landmarks (68 extracted from DLib [104]) and those in the central face", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 276 }, { "text": "regions to differentiate DeepFakes from real videos. Once these features are extracted and normalised (mean and standard deviation), a SVM is considered for the \ufb01nal classi\ufb01cation. Their proposed approach was originally evaluated with the UADFV database, achieving a \ufb01nal 89.0% AUC. However, this pre-trained model (using UADFV database) seems not to", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 277 }, { "text": "generalise very well to other databases as depicted in Table IV. Another interesting approach in this line was proposed by Agarwal and Farid in [91]. They proposed a detection system based on both facial expressions and head movements. For the feature extraction, the OpenFace2 toolkit was con- sidered [105], obtaining an intensity and occurrence for 18", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 278 }, { "text": "different facial action units related to movements of facial muscles such as cheek raiser, nose wrinkle, mouth stretch, etc. Additionally, four features related to head movements were considered. As a result, each 10-second video clip is reduced to a feature vector of dimension 190 using the Pearson correlation", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 279 }, { "text": "to measure the linearity between features. Finally, the authors considered a SVM for the \ufb01nal classi\ufb01cation. Regarding the experimental framework, the authors built their own database based on videos downloaded from YouTube of persons of in- terest talking in a formal setting, for example, weekly address,", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 280 }, { "text": "news interview, and public speech. In most videos the person is primarily facing towards the camera. Regarding the DeepFake videos, the authors trained one GAN per person based on faceswap-GAN18. Their proposed approach achieved a \ufb01nal 96.3% AUC as the best fake detection performance, being robust against new contexts and manipulation techniques.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 281 }, { "text": "Eye blinking [106] has also been studied to detect fake videos. In [92], the authors proposed an algorithm called DeepVision to analyse changes in the blinking patterns. Their approach was based on the fusion of Fast-HyperFace [107] and Eye-Aspect-Ratio (EAR) [108] to detect the face and obtain 18https://github.com/shaoanlu/faceswap-GAN", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 282 }, { "text": "9 TABLE IV IDENTITY SWAP: COMPARISON OF DIFFERENT STATE-OF-THE-ART DETECTION APPROACHES. THE BEST RESULTS ACHIEVED FOR EACH PUBLIC DATABASE ARE REMARKED IN BOLD. RESULTS IN italics INDICATE THAT THEY WERE PUBLISHED IN [26], BUT NOT IN THE ORIGINAL WORK. FF++ = FACEFORENSICS++, AUC = AREA UNDER THE CURVE, ACC. = ACCURACY, EER = EQUAL ERROR RATE, TCR = TRUE CLASSIFICATION RATES.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 283 }, { "text": "Study Method Classi\ufb01ers Best Performance Databases Korshunov and Marcel (2018) [1] Audio-Visual Features PCA+RNN PCA+LDA, SVM EER = 3.3% EER = 8.9% DeepfakeTIMIT (LQ) DeepfakeTIMIT (HQ) Matern et al. (2019) [89] Visual Features Logistic Regression MLP AUC = 85.1% Own AUC = 70.2% UADFV AUC = 77.0% AUC = 77.3%", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 284 }, { "text": "DeepfakeTIMIT (LQ) DeepfakeTIMIT (HQ) AUC = 78.0% FF++ / DFD AUC = 66.2% DFDC Preview AUC = 55.1% Celeb-DF Yang et al. (2019) [90] Head Pose Features SVM AUC = 89.0% UADFV AUC = 55.1% AUC = 53.2% DeepfakeTIMIT (LQ) DeepfakeTIMIT (HQ) AUC = 47.3% FF++ / DFD AUC = 55.9% DFDC Preview AUC = 54.6% Celeb-DF", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 285 }, { "text": "Agarwal and Farid (2019) [91] Head Pose and Facial Features SVM AUC = 96.3% Own (FaceSwap, HQ) Jung et al. (2020) [92] Eye Blinking Distance Acc. = 87.5% Own Li et al. (2019) [26], [93] Face Warping Features CNN AUC = 97.7% UADFV AUC = 99.9% AUC = 99.7% DeepfakeTIMIT (LQ) DeepfakeTIMIT (HQ) AUC = 93.0%", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 286 }, { "text": "FF++ / DFD AUC = 75.5% DFDC Preview AUC = 64.6% Celeb-DF Afchar et al. (2018) [94] Mesoscopic Features CNN Acc. = 98.4% Own AUC = 84.3% UADFV AUC = 87.8% AUC = 68.4% DeepfakeTIMIT (LQ) DeepfakeTIMIT (HQ) Acc. \u224390.0% Acc. \u224394.0% Acc. \u224398.0% FF++ (DeepFake, LQ) FF++ (DeepFake, HQ) FF++ (DeepFake, RAW)", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 287 }, { "text": "Acc. \u224383.0% Acc. \u224393.0% Acc. \u224396.0% FF++ (FaceSwap, LQ) FF++ (FaceSwap, HQ) FF++ (FaceSwap, RAW) AUC = 75.3% DFDC Preview AUC = 54.8% Celeb-DF Zhou et al. (2018) [95] Steganalysis Features + Deep Learning Features CNN SVM AUC = 85.1% UADFV AUC = 83.5% AUC = 73.5% DeepfakeTIMIT (LQ) DeepfakeTIMIT (HQ)", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 288 }, { "text": "AUC = 70.1% FF++ / DFD AUC = 61.4% DFDC Preview AUC = 53.8% Celeb-DF R\u00a8ossler et al. (2019) [12] Mesoscopic Features Steganalysis Features Deep Learning Features CNN Acc. \u224394.0% Acc. \u224398.0% Acc. \u2243100.0% FF++ (DeepFake, LQ) FF++ (DeepFake, HQ) FF++ (DeepFake, RAW) Acc. \u224393.0% Acc. \u224397.0% Acc. \u224399.0%", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 289 }, { "text": "FF++ (FaceSwap, LQ) FF++ (FaceSwap, HQ) FF++ (FaceSwap, RAW) Nguyen et al. (2019) [96] Deep Learning Features AE + Multi-Task Learning AUC = 65.8% UADFV AUC = 62.2% AUC = 55.3% DeepfakeTIMIT (LQ) DeepfakeTIMIT (HQ) AUC = 76.3% FF++ / DFD EER = 15.1% FF++ (FaceSwap, HQ) AUC = 53.6% DFDC Preview AUC = 54.3%", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 290 }, { "text": "Celeb-DF Nguyen et al. (2019) [97] Deep Learning Features Capsule Networks AUC = 61.3% UADFV AUC = 78.4% AUC = 74.4% DeepfakeTIMIT (LQ) DeepfakeTIMIT (HQ) AUC = 96.6% FF++ / DFD AUC = 53.3% DFDC Preview AUC = 57.5% Celeb-DF Dang et al. (2019) [17] Deep Learning Features CNN + Attention Mechanism AUC = 99.4%", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 291 }, { "text": "EER = 3.1% DFFD Dolhansky et al. (2019) [83] Deep Learning Features CNN Precision = 93.0% Recall = 8.4% DFDC Preview Wang and Dantcheva (2020) [98] Deep Learning Features 3DCNN TCR = 95.13% TCR = 92.25% FF++ (DeepFake, LQ) FF++ (FaceSwap, LQ) G\u00a8uera and Delp (2018) [99] Image + Temporal Features CNN + RNN", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 292 }, { "text": "10 the eye aspect ratio. Finally, features based on blinking count and period were extracted to decide whether the video is real or fake. This approach achieved a \ufb01nal 87.5% accuracy over a proprietary database. Another interesting research line is based on the detection of the artifacts included by the face manipulation pipeline.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 293 }, { "text": "In [93], Li and Lyu hypothesised that some DeepFake algo- rithms can only create images of limited resolution, which need to be further warped to match the original faces in the source video. Such transforms leave distinctive artifacts in the resulting DeepFake videos. Thus, the authors proposed a detection system based on CNNs in order to detect the", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 294 }, { "text": "presence of such artifacts from the detected face regions and the surrounding areas. Four different CNN models were trained from scratch: VGG16 [109], ResNet50, ResNet101, and ResNet152 [110]. Their proposed detection approach was tested using the UADFV and DeepfakeTIMIT databases, outperforming the state of the art for those databases.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 295 }, { "text": "Li et al. proposed later on in [26] an improved version of the work presented in [93]. In this case, the authors included a new spatial pyramid pooling module to better handle the variations in the resolution [111]. This detection approach was evaluated using different databases, achieving state-of-the-art results in", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 296 }, { "text": "some of them. Approaches based on mesoscopic and steganalysis features have also been proposed in the literature. Afchar et al. pro- posed in [94] two different networks composed of few layers in order to focus on the mesoscopic properties of the images: i) a CNN network comprised of 4 convolutional layers followed", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 297 }, { "text": "by a fully-connected layer (Meso-4), and ii) a modi\ufb01cation of Meso-4 consisted of a variant of the Inception module introduced in [112], named MesoInception-4. Their proposed approach was originally tested against DeepFakes using a private database, achieving a 98.4% of fake detection accuracy for the best performance. That pre-trained detection model was", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 298 }, { "text": "tested against unseen databases in [26], proving to be a robust approach in some cases such as with FaceForensics++. Zhou et al. proposed a two-stream network for face ma- nipulation detection. In particular, the authors considered a fusion of two streams: i) a face classi\ufb01cation stream based on the CNN GoogLeNet [112] to detect whether a face image is", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 299 }, { "text": "fake or not, and ii) a path triplet stream that is trained using steganalysis features of images patches with a triplet loss, and a SVM for the classi\ufb01cation. The initial system was trained to detect expression swap manipulations. Nevertheless, Li et al. evaluated in [26] the generalisation capacity of the pre-trained", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 300 }, { "text": "model (trained using SwapMe app) to detect identity swap manipulations, resulting to be one the most robust approaches against the recent Celeb-DF database [26]. An exhaustive analysis of different fake detection methods was carried out by R\u00a8ossler et al. using FaceForensics++ database [12]. Five different detection systems were eval-", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 301 }, { "text": "uated: i) a CNN-based system trained through handcrafted steganalysis features [113], ii) a CNN-based system whose convolution layers are speci\ufb01cally designed to suppress the high-level content of the image [114], iii) a CNN-based system with a global pooling layer that computes four statistics (mean, variance, maximum, and minimum) [115], iv) the", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 302 }, { "text": "CNN MesoInception-4 detection system described in [94], and \ufb01nally v) the CNN-based system XceptionNet [116] pre- trained using ImageNet database [117] and re-trained for the face manipulation detection task. In general, the detection system based on XceptionNet architecture provided the best results in both types of manipulation methods, DeepFakes and", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 303 }, { "text": "FaceSwap. In addition, the detection systems were evaluated considering different video quality levels in order to simulate the video processing of many social networks. In this real scenario, the accuracy of all detection systems decreased when lowering the video quality, remarking how challenging is this", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 304 }, { "text": "task in real scenarios. Recent deep learning methods considered in computer vi- sion have been applied to further improve the detection of identity swap manipulations. In [96], Nguyen et al. proposed a CNN system that uses multi-task learning to simultane- ously detect fake videos and locate the manipulated regions.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 305 }, { "text": "They considered a detection system based on an autoencoder. Concretely, they proposed to use a Y-shaped decoder in order to share valuable information between the classi\ufb01cation, segmentation, and reconstruction tasks, improving the overall performance by reducing the loss. Their proposed approach was evaluated with the FaceSwap manipulation method for the", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 306 }, { "text": "FaceForensics++ database [87], achieving a best performance of 15.07% EER, far from other detection approaches. In addition, this model seems not to generalise very well for other databases, with results below 80% AUC. Later on, the same authors presented in [97] a new fake detection system based on the recent Capsule Networks.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 307 }, { "text": "This approach uses fewer parameters than traditional CNN with similar performance [118]\u2013[120]. The proposed detec- tion system was originally evaluated using FaceForensics++ database with accuracies higher than 90%. The same pre- trained detection model was tested against unseen databases in [26], showing poor generalisation results, as it happens in", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 308 }, { "text": "most fake detection systems. Attention mechanisms have also been applied to further improve the training process of the detection systems. Dang et al. performed in [17] a thorough analysis of different face ma- nipulations. They proposed a detection system based on CNN and attention mechanisms to process and improve the feature", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 309 }, { "text": "maps of the classi\ufb01er model. Their proposed attention map can be implemented easily and inserted into existing backbone networks, through the inclusion of a single convolution layer, its associated loss functions, and masking the subsequent high- dimensional features. Their proposed detection approach was", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 310 }, { "text": "tested with the DFFD database (based on a combination of the previous FaceForensics++ databases and a collection of videos from the Internet). In particular, for identity swap detection, their proposed approach achieved an AUC of 99.43% and EER of 3.1%. Despite of the fact that it is dif\ufb01cult to provide a fair", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 311 }, { "text": "comparison among studies as different experimental protocols are considered, it is clear that their detection approach provides state-of-the-art results. In [83], in addition to the description of the DFDC database, the authors provided baseline results using three simple detec- tion systems: i) a small CNN model composed of 6 convo-", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 312 }, { "text": "11 image manipulations, ii) an XceptionNet model trained using only face images, and iii) an XceptionNet model trained using the full image. The detection system based on XceptionNet, considering only the face image (not the full image), provided the best results with 93.0% precision and 8.4% recall.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 313 }, { "text": "Deep learning approaches based on 3DCNN were studied in [121] in order to consider both spatial and motion informa- tion. In particular, the authors proposed fake detectors based on I3D [122] and 3D ResNet [123] approaches, achieving promis- ing results on the low quality videos of FaceForensics++.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 314 }, { "text": "Detection systems based not only on features at image level, but also at temporal level, along the frames of the video, have also been studied in the literature. G\u00a8uera and Delp proposed in [99] a temporal-aware pipeline to automatically detect fake videos. They considered a combination of CNNs and RNNs.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 315 }, { "text": "For the CNN, the authors used InceptionV3 [124] pre-trained using ImageNet database [117]. For the RNN system, they considered a LSTM model composed of one hidden layer with 2048 memory blocks. Finally, two fully-connected layers were included, providing the probabilities of the frame sequence being either real or fake. Their approach was evaluated using", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 316 }, { "text": "a proprietary database with a \ufb01nal 97.1% accuracy. In this line, Sabir et al. proposed a method to detect fake videos based on using the temporal information present in the stream [98]. The intuition behind this model is to exploit temporal discrepancies across frames. Thus, they considered a recurrent convolutional network similar to [99], trained in this", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 317 }, { "text": "study end-to-end instead of using a pre-trained model. Their proposed detection approach was tested through FaceForen- sics++ database, achieving AUC results of 96.9% and 96.3% for the DeepFake and FaceSwap methods, respectively. Only the low-quality videos were considered in the analysis. Finally, the discriminative power of each facial region for", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 318 }, { "text": "the detection of fake videos was studied in [100]. The authors considered a fake detection system based on XceptionNet. Databases from both 1st and 2nd generations were considered in the experimental framework, concluding that poor fake detection results are achieved in the latest DeepFake video databases of the 2nd generation compared with the 1st gener-", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 319 }, { "text": "ation, with results of 91.0% and 83.6% AUC for the DFDC Preview and Celeb-DF databases, respectively. It is important to highlight that, contrary to [26], a separate fake detection system was speci\ufb01cally trained for each database. In conclusion, although many different approaches have been proposed in the literature, they all show poor generalisa-", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 320 }, { "text": "tion results to unseen databases, as indicated in Table IV. In addition, we also highlight the poor detection results achieved by most approaches on the DeepFake databases of the 2nd generation with results below 60% AUC. V. ATTRIBUTE MANIPULATION A. Manipulation Techniques and Public Databases This face manipulation consists of modifying in an image", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 321 }, { "text": "some attributes of the face such as the colour of the hair or the skin, the gender, the age, adding glasses, etc. Despite the success of GAN-based frameworks for general image translations and manipulations [43], [71], [125]\u2013[129], and in particular for face attribute manipulations [43], [44], [68],", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 322 }, { "text": "GAN approaches in the \ufb01eld, from older to closer in time, providing also the link to their corresponding codes. In [130], the authors introduced the Invertible Conditional GAN (IcGAN)19 for complex image editing as the union of an encoder used jointly with a conditional GAN (cGAN) [135]. This approach provides accurate results in terms of attribute", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 323 }, { "text": "manipulation. However, it seriously changes the face identity of the person. Lample et al. proposed in [133] an encoder-decoder archi- tecture that is trained to reconstruct images by disentangling the salient information of the image and the attribute values directly in the latent space20. However, as it happens with the", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 324 }, { "text": "IcGAN approach, the generated images may lack some details or present unexpected distortions. An enhanced approach named StarGAN21 was proposed in [43]. Before the StarGAN approach, many studies had shown promising results in image-to-image translations for two domains in general. However, few studies had focused on", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 325 }, { "text": "handling more than two domains. In that case a direct approach would be to build different models independently for every pair of image domains. StarGAN proposed a novel approach able to perform image-to-image translations for multiple domains using only a single model. The authors trained a conditional", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 326 }, { "text": "attribute transfer network via attribute classi\ufb01cation loss and cycle consistency loss. Good visual results were achieved compared with previous approaches. However, it sometimes includes undesired modi\ufb01cations from the input face image such as the colour of the skin. Almost at the same time He et al. proposed in [68]", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 327 }, { "text": "attGAN22, a novel approach that removes the strict attribute- independent constraint from the latent representation, and just applies the attribute-classi\ufb01cation constraint to the generated image to guarantee the correct change of the attributes. AttGAN provides state-of-the-art results on realistic attribute", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 328 }, { "text": "manipulation with other facial details well preserved. One of the latest approaches proposed in the literature is STGAN23 [44]. In general, attribute manipulation can be tack- led by incorporating an encoder-decoder or GAN. However, as commented Liu et al. [44], the bottleneck layer in the encoder-decoder usually provides blurry and low quality ma-", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 329 }, { "text": "nipulation results. To improve this, the authors presented and incorporated selective transfer units with an encoder-decoder for simultaneously improving the attribute manipulation ability and the image quality. As a result, STGAN has recently outperformed the state of the art in attribute manipulation.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 330 }, { "text": "Despite of the fact that the code of most attribute ma- nipulation approaches are publicly available, the lack of 19https://github.com/Guim3/IcGAN 20https://github.com/facebookresearch/FaderNetworks 21https://github.com/yunjey/stargan/blob/master/README.md 22https://github.com/LynnHo/AttGAN-Tensor\ufb02ow", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 331 }, { "text": "12 TABLE V ATTRIBUTE MANIPULATION: COMPARISON OF DIFFERENT STATE-OF-THE-ART DETECTION APPROACHES. THE BEST RESULTS ACHIEVED FOR EACH PUBLIC DATABASE ARE REMARKED IN BOLD. AUC = AREA UNDER THE CURVE, ACC. = ACCURACY, EER = EQUAL ERROR RATE. Study Method Classi\ufb01ers Best Performance Databases (Generation)", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 332 }, { "text": "Wang et al. (2019) [56] GAN-Pipeline Features SVM Acc. = 84.7% Own (InterFaceGAN/StyleGAN) Nataraj et al. (2019) [58] Steganalysis Features CNN Acc. = 99.4% Own (StarGAN/CycleGAN) Bharati et al. (2016) [136] Deep Learning Features (Face Patches) RBM Overall Acc. = 96.2% Overall Acc. = 87.1% Own (Celebrity Retouching,", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 333 }, { "text": "ND-IIITD Retouching) Jain et al. (2019) [137] Deep Learning Features (Face Patches) CNN + SVM Overall Acc. = 99.6% Overall Acc. = 99.7% Own (ND-IIITD Retouching, StarGAN) Tariq et al. (2018) [138] Deep Learning Features CNN AUC = 99.9% AUC = 74.9% Own (ProGAN, Adobe Photoshop) Dang et al. (2019) [17]", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 334 }, { "text": "Deep Learning Features CNN + Attention Mechanism AUC = 99.9% EER = 1.0% DFFD (FaceApp/StarGAN) Wang et al. (2019) [139] Deep Learning Features DRN AP = 99.8% Own (Adobe Photoshop) Marra et al. (2019) [60] Deep Learning Features CNN + Incremental Learning Acc. = 99.3% Own (Glow/StarGAN ) Zhang et al. (2019)", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 335 }, { "text": "comparison among studies. Up to now, to the best of our knowledge, the DFFD database [17] seems to be the only public database that considers this type of facial manipulations. This database comprises 18,416 and 79,960 fake images gener- ated through FaceApp and StarGAN approaches, respectively. B. Manipulation Detection", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 336 }, { "text": "Attribute manipulations have been originally studied in the \ufb01eld of face recognition in order to see how robust biometric systems are against physical factors such as plastic surgery, cosmetics, makeup or occlusions [142]\u2013[147]. However, it has been the recent success of mobile applications such as FaceApp that has motivated the research community to", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 337 }, { "text": "detect digital face attribute manipulations. Table V provides a comparison of the most relevant approaches in this area. We include for each study information related to the method, classi\ufb01ers, best performance, and databases for research. Some authors propose to analyse the internal GAN pipeline to detect different artifacts between real and manipulated", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 338 }, { "text": "images. Similar to the entire face synthesis manipulations, Wang et al. conjectured in [56] that monitoring neuron be- havior could also serve as an asset in detecting fake faces since layer-by-layer neuron activation patterns may capture more subtle features that are important for the facial ma- nipulation detection system. Their proposed approach, named", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 339 }, { "text": "FakeSpoter, extracted as features neuron coverage behaviors of real and fake faces from deep face recognition systems (VGG-Face [63], OpenFace [64], and FaceNet [65]), and then trained a SVM for the \ufb01nal classi\ufb01cation. The authors tested their proposed approach using real faces from CelebA-HQ [48] and FFHQ [41] databases and synthetic faces created through", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 340 }, { "text": "InterFaceGAN [66] and StyleGAN [41], achieving for the best performance a \ufb01nal 84.7% manipulation detection accuracy using the FaceNet model. Fake detection systems inspired in steganalysis have also been studied. As described in Sec. III-B for the entire face synthesis, Nataraj et al. proposed in [58] a detection system", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 341 }, { "text": "based on the combination of pixel co-occurrence matrices and CNN. They created a new fake dataset based on attribute ma- nipulations using the StarGAN approach [43] trained through the CelebA database [52], achieving a \ufb01nal 99.4% accuracy for the best result. Many studies have also focused on pure deep learning", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 342 }, { "text": "methods, either feeding the networks with face patches or with the complete face. In [136], Bharati et al. proposed a deep learning approach based on a Restricted Boltzmann Machine (RBM) in order to detect digital retouching of face images. The input of the detection system consisted of face patches in order to learn discriminative features to classify", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 343 }, { "text": "each image as original or retouched. Regarding the databases, the authors generated two fake databases from the original ND-IIITD database (collection B [148]) and a set of celebrity facial images downloaded from the Internet. Fake images were generated using the professional software PortraitPro Studio", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 344 }, { "text": "Max24, considering aspects such as skin texture, shape of eyes, nose, lips and overall face, prominence of smile, lip shape, and eye colour. Their proposed approach achieved overall accuracies for manipulation detection of 96.2% and 87.1% for the celebrity and ND-IIITD retouching databases, respectively.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 345 }, { "text": "13 A similar approach based on non-overlapping face patches was presented in [137]. Jain et al. proposed a CNN feature extractor composed of 6 convolutional layers and 2 fully- connected layers. Additionally, residual connections were con- sidered inspired by a ResNet architecture [110]. Finally, a", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 346 }, { "text": "SVM was used for the \ufb01nal classi\ufb01cation. Regarding the experimental framework, the ND-IIITD retouched database presented in [136] was considered. Additionally, the authors considered fake images created through the StarGAN ap- proach [43], trained using the CelebA database [52]. In gen- eral, good detection results were achieved in both manipulation", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 347 }, { "text": "approaches, achieving almost 100% manipulation detection accuracy. Deep learning methods based on the complete face have been further studied in the literature, achieving in general very good results. Tariq et al. evaluated in [138] the use of differ- ent CNN architectures such as VGG16 [63], VGG19 [63],", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 348 }, { "text": "ResNet [110], or XceptionNet [116], among others. For the real face images, the CelebA database [52] was used. Re- garding the fake images, two different approaches were con- sidered: i) machine approaches based on GAN, in particular ProGAN [48], and ii) manual approach based on Adobe Photo- shop CS6, including manipulations such as makeup, glasses,", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 349 }, { "text": "sunglasses, hair, and hats. For the experimental evaluation, different sizes of the images were considered (from 32\u00d732 to 256\u00d7256 pixels). A \ufb01nal 99.99% AUC was obtained for the machine-created scenario whereas for the human-created scenario this value decreased to a \ufb01nal 74.9% AUC for the best CNN model. Thus, a high degradation of the manipulation", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 350 }, { "text": "detection performance was observed between machine- and human-created fake images. Attention mechanisms have also been applied to further improve the training process of the detection systems. As described in previous sections, Dang et al. developed in [17] a system able to detect different types of fakes. They used", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 351 }, { "text": "attention mechanisms to process and improve the feature maps of CNN models. Regarding the attribute manipulations, two different approaches were considered: i) fake images created through the public FaceApp software, with up to 28 different available \ufb01lters considering aspects such as hair, age, glasses,", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 352 }, { "text": "beard, and skin colour, among others; and ii) fake images created through the StarGAN approach [43], with up to 40 different \ufb01lters. Their proposed approach was tested using their novel database DFFD, achieving very good results close to 1.0% EER (and 99.9% of AUC). Wang et al. carried out in [139] an interesting research", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 353 }, { "text": "using publicly available commercial software from Adobe Photoshop (Face-Aware Liquify tool [149]) in order to syn- thesise new faces, and also a professional artist in order to manipulate 50 real photographs. The authors began running a human study through Amazon Mechanical Turk (AMT), showing real and fake images to the participants and asking", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 354 }, { "text": "them to classify each image into one of the classes. The results achieved remark how challenging the task is for humans, with a \ufb01nal 53.5% of accuracy, close to chance (50%). After the human study, the authors proposed two different automatic models: i) a global classi\ufb01cation model based on Dilated", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 355 }, { "text": "Residual Networks (DRN) to predict whether the face has been warped or not, and ii) a local warp predictor based on the optical \ufb02ow \ufb01eld in order to identify where manipulation occurs, and reverse them. The PWC-Net approach proposed in [150] was considered to compute the \ufb02ow from original to manipulated and vice versa. Performances of 99.8% and", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 356 }, { "text": "97.4% for automatic and manual face synthesis manipulation were achieved. The work [60] by Marra et al. also described in Sec. III-B was able to correctly perform discrimination when new GANs were presented to the network and achieved a 99.3% accuracy for their proposed manipulation detection approach, based on", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 357 }, { "text": "the XceptionNet model. A detection system based on features extracted from the spectrum domain, rather than the raw image pixels, was presented by Zhang et al. in [140]. Given an image as input, they applied a 2D DFT to each of the RGB channels, getting one frequency image per channel. Regarding the classi\ufb01er,", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 358 }, { "text": "they proposed AutoGAN, which is a GAN simulator that can synthesise GAN artifacts in any image without needing to access any pre-trained GAN model. The generalisation capacity of their proposed approach was tested using unseen GAN models. In particular, StarGAN [43] and GauGAN [126] were considered in the evaluation. For the StarGAN approach,", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 359 }, { "text": "good detection results were achieved using the frequency domain (100%). However, for the GauGAN approach, a high degradation of the system performance, 50% accuracy, was observed. The authors claimed that this was produced due to the generator of the GauGAN is drastically different from the CycleGAN (used in training).", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 360 }, { "text": "Finally, Rathgeb et al. proposed in [141] a detection system based on Photo Response Non-Uniformity (PRNU). Speci\ufb01- cally, scores obtained from the analysis of spatial and spectral features extracted from PRNU patterns across image cells were fused. Their proposed approach was evaluated over a private database created using 5 different mobile applications,", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 361 }, { "text": "achieving an average 13.7% EER in manipulation detection. To summarise this section, we can see that the core of most attribute manipulation detection systems are based on deep learning technology, providing in general very good results close to 100% accuracy, as indicated in Table V. This is mainly produced due to the GAN-\ufb01ngerprint information", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 362 }, { "text": "present in fake images. However, as indicated in the entire face synthesis manipulation, recent studies have been proposed in the literature to remove such GAN \ufb01ngerprints from the fake images while keeping very realistic appearance [16], [78], which represent a challenge even for the most advanced", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 363 }, { "text": "manipulation detectors. VI. EXPRESSION SWAP A. Manipulation Techniques and Public Databases This manipulation, also known as face reenactment, consists of modifying the facial expression of the person. We focus on the most popular techniques Face2Face and NeuralTextures, which replace the facial expression of one person in a video", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 364 }, { "text": "14 Initially, the FaceForensics database was focused on the Face2Face approach [45]. This is a computer graphics ap- proach that transfers the expression of a source video to a target video while maintaining the identity of the target person. This was carried out through manual keyframe selection. Concretely, the \ufb01rst frames of each video were used to obtain", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 365 }, { "text": "a temporary face identity (i.e., a 3D model), and track the expression over the remaining frames. Then, fake videos were generated by transferring the source expression parameters of each frame (i.e., 76 Blendshape coef\ufb01cients) to the target video. Later on, the same authors presented in FaceForen-", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 366 }, { "text": "sics++ a new learning approach based on NeuralTextures [46]. This is a rendering approach that uses the original video data to learn a neural texture of the target person, including a rendering network. In particular, the authors considered in their implementation a patch-based GAN-loss as used in Pix2Pix [126]. Only the facial expression corresponding to", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 367 }, { "text": "the mouth was modi\ufb01ed. It is important to remark that all data is available on the FaceForensics++ GitHub25. In total, there are 1,000 real videos extracted from Youtube. Regarding the manipulated videos, 2,000 fake videos are available (1,000 videos for each considered fake approach). In addition, it is", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 368 }, { "text": "important to highlight that different video quality levels are considered, in particular: i) RAW (original quality), ii) HQ (constant rate quantization parameter equal to 23), and iii) LQ (constant rate quantization parameter equal to 40). This aspect simulates the video processing techniques usually applied in", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 369 }, { "text": "social networks. In addition to the Face2Face and NeuralTexture techniques considered in expression swap manipulations at video level, different approaches have been recently proposed to change the facial expression in both images and videos. A very popular approach was presented in [151]. Averbuch-Elor et", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 370 }, { "text": "al. proposed a technique to automatically animate a still portrait using a video of a different subject, transferring the expressiveness of the subject of the video to the target portrait. Unlike Face2Face and NeuralTexture approaches that require videos from both input and target faces, in [151] just an image", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 371 }, { "text": "of the target is needed. In this line, a recent approach was recently presented in [152], providing very good results in both one-shot and few-shot learning. Finally, we also highlight other popular approaches at image level. For example, mobile applications such as FaceApp26 allow to easily change the level of smiling, from happier", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 372 }, { "text": "to angrier. These approaches are based on current GAN architectures. For example, Choi et al. showed in [43] the potential of StarGAN to change an input image to different expression levels such as angry, happy, neutral, sad, surprised, and fearful. Other recent GAN approaches that improve both the image quality of the fake images and the control editing", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 373 }, { "text": "of the parameters are InterFaceGAN [66], UGAN [153], STGAN [44], and AttGAN [68]. 25https://github.com/ondyari/FaceForensics 26https://apps.apple.com/gb/app/faceapp-ai-face-editor/id1180884341 B. Manipulation Detection This section aims to provide an overview of the expression swap detectors at video level using the FaceForensics++", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 374 }, { "text": "database, as this is the only publicly available database for research in this area, to the best of our knowledge. Manipu- lations at image level (not video) can be detected using the same approaches described in Sec. III-B and V-B. Table VI provides a comparison of the most relevant approaches in the area of expression swap detection. For", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 375 }, { "text": "each study we include information related to the method, classi\ufb01ers, best performance, and databases. We highlight in bold the best results achieved for the only public database, FaceForensics++. It is important to remark that in some cases, different evaluation metrics are considered (e.g., AUC and", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 376 }, { "text": "EER), which makes it dif\ufb01cult to perform a fair comparison among the studies. Some of the following methods were already discussed in Sect. IV-B for identity swap. Here we summarise the results achieved by them in detecting expression swap manipulations. Preliminary studies have focused on the visual features ex-", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 377 }, { "text": "isted in fake videos such as the eye colour, missing re\ufb02ections, etc. In [89] by Matern et al., the proposed approach was tested using FaceForensics++, only the Face2Face manipulation tech- nique, achieving a \ufb01nal 86.6% AUC for the best performance. Approaches based on mesoscopic and steganalysis features", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 378 }, { "text": "have also been studied in the literature. In [94], the proposed approach was tested using the Face2Face fake videos from the FaceForensics++ database [12], achieving in general good results, especially for RAW-quality videos. The same approach was later on tested in [12] against NeuralTextures fake videos,", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 379 }, { "text": "obtaining lower accuracy results compared with Face2Face. Recent deep learning methods have also been applied with good results. In [12], the detection system based on XceptionNet provided the best results in both Face2Face and NeuralTextures manipulations, close to 100% on RAW quality. In addition, the detection systems were evaluated considering", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 380 }, { "text": "different video quality levels in order to simulate the video processing of many social networks. In this real scenario, the accuracy of all detection systems was degraded with the video quality, as it happens in identity swap manipulations. In [96], the proposed approach based on multi-task learning", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 381 }, { "text": "was evaluated with the FaceForensics++ database. For the Face2Face method, a 7.1% EER was achieved on HQ videos whereas for the NeuralTexture method, the EER increased a bit more to a \ufb01nal 7.8% EER in manipulation detection. Attention mechanisms have been recently proposed in [17] to further improve the training process. The proposed detection", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 382 }, { "text": "approach was tested using the DFFD database, which for the expression swap manipulation is based only on data from FaceForensics++ database. The proposed approach achieved an AUC = 99.4% and EER = 3.4%. Deep learning approaches based on 3DCNN were stud- ied in [121] in order to consider both spatial and motion", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 383 }, { "text": "15 TABLE VI EXPRESSION SWAP: COMPARISON OF DIFFERENT STATE-OF-THE-ART DETECTION APPROACHES. THE BEST RESULTS ACHIEVED FOR EACH PUBLIC DATABASE ARE REMARKED IN BOLD. FF++ = FACEFORENSICS++, AUC = AREA UNDER THE CURVE, ACC. = ACCURACY, EER = EQUAL ERROR RATE, TCR = TRUE CLASSIFICATION RATE. Study Method", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 384 }, { "text": "Classi\ufb01ers Best Performance Databases (Generation) Matern et al. (2019) [89] Visual Features Logistic Regression, MLP AUC = 86.6% FF++ (Face2Face, RAW) Afchar et al. (2018) [94] Mesoscopic Features CNN Acc. = 83.2% FF++ (Face2Face, LQ) Acc. = 93.4% FF++ (Face2Face, HQ) Acc. = 96.8% FF++ (Face2Face, RAW)", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 385 }, { "text": "Acc. \u224375% FF++ (NeuralTextures, LQ) Acc. \u224385% FF++ (NeuralTextures, HQ) Acc. \u224395% FF++ (NeuralTextures, RAW) R\u00a8ossler et al. (2019) [12] Mesoscopic Features Steganalysis Features Deep Learning Features CNN Acc. \u224391% FF++ (Face2Face, LQ) Acc. \u224398% FF++ (Face2Face, HQ) Acc. \u2243100% FF++ (Face2Face, RAW)", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 386 }, { "text": "Acc. \u224381% FF++ (NeuralTextures, LQ) Acc. \u224393% FF++ (NeuralTextures, HQ) Acc. \u224399% FF++ (NeuralTextures, RAW) Nguyen et al. (2019) [96] Deep Learning Features Autoencoder EER = 7.1% FF++ (Face2Face, HQ) EER = 7.8% FF++ (NeuralTextures, HQ) Dang et al. (2020) [17] Deep Learning Features CNN + Attention Mechanism", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 387 }, { "text": "AUC = 99.4% EER = 3.4% FF++ (Face2Face, -) Wang and Dantcheva (2020) [98] Deep Learning Features 3DCNN TCR = 90.27% TCR = 80.5% FF++ (Face2Face, LQ) FF++ (NeuralTextures, LQ) Sabir et al. (2019) [98] Image + Temporal Features CNN + RNN Acc. = 94.3 FF++ (Face2Face, LQ) Amerini et al. (2019) [154] Image + Temporal Features", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 388 }, { "text": "CNN + Optical Flow Acc. = 81.6% FF++ (Face2Face, -) Another interesting line is based on the analysis of both im- age and temporal information. In [98], the proposed approach based on recurrent convolutional networks was tested using the FaceForensics++ database, achieving AUC results of 94.3% for the Face2Face technique. Only the low-quality videos were", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 389 }, { "text": "considered in the analysis. Finally, in [154], Amerini et al. proposed the adoption of optical \ufb02ow \ufb01elds to exploit possible inter-frame dissimilarities, using the PWC-Net approach [150]. The optical \ufb02ow is a vector \ufb01eld computed among consecutive frames to extract apparent motion in the scene. The use of this", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 390 }, { "text": "approach is motivated as fake videos should have unnatural optical \ufb02ow due to the unusual movement of lips, eyes, etc. Preliminary results were obtained using both VGG16 and ResNet50 networks, obtaining an Acc. = 81.6% for the best performance in manipulation detection. Finally, as stated previously, most of the approaches re-", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 391 }, { "text": "ported here for expression swap detection have also been used for identity swap detection as reviewed in Sec. IV-B. In general, it seems that similar features can be learnt by the fake detectors to distinguish between real and fake content, achiev- ing good results in both types of manipulations. We highlight", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 392 }, { "text": "the potential of novel techniques such as attention mechanisms to better guide the networks during the training process, as shown in [17], achieving AUC results of 99.4% for detecting both identity swap and expression swap manipulations. VII. OTHER FACE MANIPULATION DIRECTIONS The four classes of face manipulation techniques described", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 393 }, { "text": "before are the ones that are receiving most attention in the last few years, but they do not perfectly represent all possible face manipulations. This section discusses some other challenging and dangerous approaches in face manipulation: face morphing, face de-identi\ufb01cation, and face synthesis based", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 394 }, { "text": "on audio or text (i.e., audio-to-video and text-to-video). A. Face Morphing Face morphing is a type of face manipulation that can be used to create arti\ufb01cial biometric face samples that resemble the biometric information of two or more individuals [155], [156]. This means that the new morphed face image would", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 395 }, { "text": "be successfully veri\ufb01ed against facial samples of these two or more individuals creating a serious threat to face recognition systems [157], [158]. In this sense, face morphing is a different type of facial manipulation compared with the four main types covered in this survey. Also, it is worth noting that face", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 396 }, { "text": "morphing is mainly focused on creating fake samples at image level, not video such as identity swap manipulations. There has been recently a large amount of research in the \ufb01eld of face morphing. A very complete review of this \ufb01eld has been published by Scherhag et al. [156] in 2019 including both morphing techniques and also morphing attack detectors.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 397 }, { "text": "Despite the large amount of publications, the research in this \ufb01eld is still in its infancy, with many open issues and challenges. It is important to highlight the lack of publicly available databases and benchmarks what makes it dif\ufb01cult to perform a fair comparison among studies. In order to overcome", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 398 }, { "text": "16 mark27. The database comprises morphed and real images constituting 1,800 photographs of 150 subjects. Morphing images were generated using 6 different algorithms, presenting a wide variety of possible approaches. Regarding the face morphing detectors, different approaches have been proposed in the literature based on different features,", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 399 }, { "text": "e.g.: the reduction of face details due to blending opera- tions [160], Fourier spectrum of sensor pattern noise [161], dif- ferences between the facial landmarks [162], [163], and pure deep learning features [164], [165]. In addition, approaches based on face de-morphing have been studied in order to", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 400 }, { "text": "restore the accomplice\u2019s facial image [166], [167]. B. Face De-Identi\ufb01cation The main goal of face de-identi\ufb01cation (de-ID) is to remove the identity information present on a face image or video in order to preserve the privacy of the person [168]. This can be achieved in several ways. The simplest way can be", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 401 }, { "text": "just to obfuscate the face by blurring or pixelation (e.g., in Google Maps Street View). More sophisticated methods try to provide face images with different identities but maintaining all other factors (pose, expression, illumination, etc.) unaltered. Therefore, the concept of face de-ID is very general. One", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 402 }, { "text": "possible option to achieve face de-identi\ufb01cation could be through face identity swap. Earlier works in this area were based on applying face de- ID to still images. In [169] Gross et al. presented a multi- factor framework for de-ID, which combined linear, bilinear, and quadratic models. They showed their method was able to", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 403 }, { "text": "protect privacy while preserving data utility on an expression- variant face database. Recently, the developments of image synthesis methods based on generative deep neural networks, in particular GAN, have inspired new face de-ID methods such as [170]\u2013[175], which use synthesised faces to replace the", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 404 }, { "text": "original ones. Also, in [176], the authors proposed the use of Semi-Adversarial Networks (SAN) to confound arbitrary face-based gender classi\ufb01ers. More recently, in [177] Gafni et al. presented in 2019 a method that provides face de-ID with convincing performance even in unconstrained videos. Their approach is based on an", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 405 }, { "text": "adversarial autoencoder coupled with a trained face classi\ufb01er. This way they can achieve a rich latent space, embedding both identity and expression information. Also, in [178] a new face de-ID method based on a deep transfer model was presented. This method treats the non-identity related facial attributes as", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 406 }, { "text": "the style of the original faces, and uses a trained facial attribute transfer model to extract and map them to different faces achieving very promising results both in images and videos. Some other related studies in this area work directly over face representations or deep face models by eliminating there", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 407 }, { "text": "undesired or protected information like identity, gender, or facial expressions [179]\u2013[181]. Once that protected informa- tion has been disentangled, a face image or video can then be generated based on the new representations originated in which the protected information has been eliminated, reduced,", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 408 }, { "text": "or obfuscated. 27https://biolab.csr.unibo.it/fvcongoing C. Audio-to-Video and Text-to-Video A related topic to facial expression swap is the synthesis of video from audio or text. These types of video face manipulations are also known as lip-sinc deep fakes [182]. Popular examples can be seen on the Internet28 29.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 409 }, { "text": "Regarding the synthesis of fake videos from audio (audio- to-video), Suwajanakorn et al. presented in [125] an approach to synthesise high quality videos of a person (Obama in this case) speaking with accurate lip sync. For this, they used as input to their approach many hours of previous videos of the person together with a new audio recording. In their", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 410 }, { "text": "approach they employed a recurrent neural network (based on LSTMs) to learn the mapping from raw audio features to mouth shapes. Then, based on the mouth shape at each frame, they synthesised high quality mouth texture, and composited it with 3D pose matching to create the new video to match the input audio track, producing photorealistic results.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 411 }, { "text": "In [183], Song et al. proposed an approach based on a novel conditional recurrent generation network that incorporates both image and audio features in the recurrent unit for temporal dependency, and also a pair of spatial-temporal discriminators for better image/video quality. As a result, their approach can", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 412 }, { "text": "model both lip and mouth together with expression and head pose variations as a whole, achieving much more realistic results. The source code is publicly available in GitHub30. Also, in [184] Song et al. presented a dynamic method not assuming a person-speci\ufb01c rendering network like in [125]. In their approach they are able to generate very realistic fake", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 413 }, { "text": "videos by carrying out a 3D face model reconstruction from the input video plus a recurrent network to translate the source audio into expression parameters. Finally, they introduced a novel video rendering network and a dynamic programming method to construct a temporally coherent and photo-realistic", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 414 }, { "text": "video. Video results are shown on the Internet31. Another interesting approach was presented in [185]. Zhou et al. proposed a novel framework called Disentangled Audio- Visual System (DAVS), which generates high quality talking face videos using disentangled audio-visual representation. Both audio and video speech information can be employed", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 415 }, { "text": "as input guidance. The source code is available in GitHub32. Regarding the synthesis of fake videos from text (text-to- video), Fried et al. proposed in [186] a method that takes as input a video of a person speaking and the desired text to be spoken, and synthesises a new video in which the persons", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 416 }, { "text": "mouth is synchronised with the new words. In particular, their method automatically annotates an input talking-head video with phonemes, visemes, 3D face pose and geometry, re- \ufb02ectance, expression and scene illumination per frame. Finally, a recurrent video generation network creates a photorealistic", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 417 }, { "text": "video that matches the edited transcript. Examples of the fake videos generated with this approach are publicly available33. 28https://www.youtube.com/watch?v=VWMEDacz3L4 29https://www.bbc.com/news/technology-48607673 30https://github.com/susanqq/Talking Face Generation 31https://wywu.github.io/projects/EBT/EBT.html", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 418 }, { "text": "17 To the best of our knowledge, there are no publicly available databases and benchmarks related to audio- and text-to-video fake detection content. Research on this topic is usually carried out through the synthesis of in-house data using publicly available implementations like the ones described in this", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 419 }, { "text": "section. Recent studies have analysed how easy is to detect audio- and text-to-video fake content. In [182], Agarwal et al. pro- posed a fake detection method that exploits the inconsistencies that exist between the dynamics of the mouth shape (visemes) and the spoken phoneme. They focused on some particular", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 420 }, { "text": "visemes in which the mouth must be completely closed and observed that this did not happen in many manipulated videos. Their proposed approach achieved good results, specially as the length of the video increases. VIII. CONCLUDING REMARKS Motivated by the ongoing success of digital face manipu- lations, specially DeepFakes, this survey provides a compre-", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 421 }, { "text": "hensive panorama of the \ufb01eld, including details of up-to-date: i) types of facial manipulations, ii) facial manipulation tech- niques, iii) public databases for research, and iv) benchmarks for the detection of each facial manipulation group, including key results achieved by the most representative manipulation", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 422 }, { "text": "detection approaches. Generally speaking, most current face manipulations seem easy to be detected under controlled scenarios, i.e., when fake detectors are evaluated in the same conditions they are trained for. This fact has been demonstrated in most of the benchmarks included in this survey, achieving very low error", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 423 }, { "text": "rates in manipulation detection. However, this scenario may not be very realistic as fake images and videos are usually shared on social networks, suffering from high variations such as compression level, resizing, noise, etc. Also, facial manip- ulation techniques are continuously improving. These factors", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 424 }, { "text": "motivate further research on the generalisation ability of the fake detectors against unseen conditions. This aspect has been preliminary studied in different works [16], [59]\u2013[61]. Future research could be in the line of the latest publications [187], [188] as they do not require fake videos for training, providing", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 425 }, { "text": "a better generalisation ability to unseen attacks. Fusion techniques, at a feature or score level, could provide a better adaptation of the fake detectors to the different scenarios [189]\u2013[191]. In fact, different fake detection ap- proaches are already based on the combination of different sources of information, e.g., Zhou et al. proposed in [95] a", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 426 }, { "text": "detection system based on the combination of steganalysis and pure deep learning features, whereas Rathgeb et al. proposed in [141] the combination of spatial and spectral features. Another two interesting fusion approaches have been recently presented in [192], [193], combining RGB, Depth, and InfraRed information to detect physical face attacks. Also,", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 427 }, { "text": "face weighting approaches have been proposed in order to detect fake videos using multiple frames [194]. Finally, fusion of other sources of information such as the text, keystroke, or the audio that accompanies the videos when uploading them to social networks could be very valuable to improve the detectors [195]\u2013[198].", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 428 }, { "text": "In addition to the traditional fake detectors based only on the image/video information, novel schemes should be studied in order to provide more robust tools. One example of this is the work presented by Tursman et al. in [199]. The authors proposed to detect fake content via social veri\ufb01cation at capture time: the arbiters of truthfulness are a group of video", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 429 }, { "text": "cameras that synchronously capture a speaker, collectively reach consensus, and then sign their videos in real time as \u201ctrue\u201d. Approaches like this one could further protect media content from attacks. We highlight next the key aspects to improve and future trends to follow for each facial manipulation group:", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 430 }, { "text": "\u2022 Face Synthesis: current manipulations are usually based on GAN architectures such as StyleGAN, providing very realistic images. Nevertheless, most detectors can easily distinguish between real and fake images, achieving accu- racies close to 100%. This is produced due to fake images are characterised by speci\ufb01c GAN \ufb01ngerprints. But, what", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 431 }, { "text": "if we are able to remove those GAN \ufb01ngerprints or add some noise patterns while keeping very realistic synthetic images? Recent approaches have focused on this research line, which represents a challenge even for the best manipulation detection systems [16], [78], [200]. \u2022 Identity Swap: although many different approaches have", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 432 }, { "text": "been proposed in the literature, it is certainly dif\ufb01cult to decide which is the best one. This is produced due to many different factors. First, most approaches are trained for a speci\ufb01c database and compression level, achieving in general very good results. However, they all show poor generalisation results to unseen conditions. In addition,", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 433 }, { "text": "the fact that different metrics (i.e., Acc., AUC, EER, etc.) and experimental protocols are usually considered does not help to achieve fair comparisons among studies. All these aspects should be further considered to advance in the \ufb01eld. Furthermore, we want to highlight the detection results achieved in the latest DeepFake databases of the 2nd", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 434 }, { "text": "generation such as DFDC and Celeb-DF [26], [83]. While fake detectors already achieve AUC results close to 100% in databases of the 1st generation such as UADFV and FaceForensics++ [12], [82], they all suffer from a high performance degradation on the latest ones, in particular for the Celeb-DF database with AUC results below 60%", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 435 }, { "text": "in most cases. Therefore, more efforts are needed to fur- ther improve current fake detection systems, for example, through large-scale challenges and benchmarks such as the recent DFDC34. \u2022 Attribute Manipulation: the same aspect highlighted for the face synthesis (GAN \ufb01ngerprint removal) also applies here as most manipulations are based on GAN", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 436 }, { "text": "18 \u2022 Expression Swap: contrary to the identity swap, which has rapidly evolved with the release of improved Deep- Fake databases, the only public database in expression swap is FaceForensics++, to the best of our knowledge. This database is characterised by visual artifacts that are easy to detect, achieving therefore AUC results close to", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 437 }, { "text": "100% in several fake detection approaches. We encourage researchers to generate and make public more realistic databases based on recent techniques [125], [151], [184]. All these aspects, together with the development of im- proved GAN approaches and the recent DeepFake Detection Challenge (DFDC) will foster the new generation of realistic", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 438 }, { "text": "fake images/videos [70] together with more advanced tech- niques for face manipulation detection. ACKNOWLEDGMENTS This work has been supported by projects: PRIMA (H2020- MSCA-ITN-2019-860315), TRESPASS-ETN (H2020-MSCA- ITN-2019-860813), BIBECA (MINECO/FEDER RTI2018- 101248-B-I00), Bio-Guard (Ayudas", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 439 }, { "text": "Fundaci\u00b4on BBVA a Equipos de Investigaci\u00b4on Cient\u00b4\u0131\ufb01ca 2017), and Accenture. Ruben Tolosana is supported by Consejer\u00b4\u0131a de Educaci\u00b4on, Juventud y Deporte de la Comunidad de Madrid y Fondo Social Europeo. REFERENCES [1] P. Korshunov and S. Marcel, \u201cDeepfakes: a New Threat to Face Recog- nition? Assessment and Detection,\u201d arXiv preprint arXiv:1812.08685,", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 440 }, { "text": "2018. [2] D. Citron, \u201cHow DeepFake Undermine Truth and Threaten Democracy,\u201d 2019. [Online]. Available: https://www.ted.com [3] R. Cellan-Jones, \u201cDeepfake Videos Double in Nine Months,\u201d 2019. [Online]. Available: https://www.bbc.com/news/technology-49961089 [4] BBC Bitesize, \u201cDeepfakes: What Are They", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 441 }, { "text": "and Why Would I Make One?\u201d 2019. [Online]. Available: https://www.bbc.co.uk/ bitesize/articles/zfkwcqt [5] A. Swaminathan, M. Wu and K.J.R. Liu, \u201cDigital Image Forensics via Intrinsic Fingerprints,\u201d IEEE Transactions on Information Forensics and Security, vol. 3, no. 1, pp. 101\u2013117, 2008. [6] H. Farid, \u201cImage Forgery Detection,\u201d IEEE Signal Processing Maga-", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 442 }, { "text": "zine, vol. 26, no. 2, pp. 16\u201325, 2009. [7] M. Stamm and K. Liu, \u201cForensic Detection of Image Manipulation Using Statistical Intrinsic Fingerprints,\u201d IEEE Transactions on Infor- mation Forensics and Security, vol. 5, no. 3, pp. 492\u2013506, 2010. [8] A. Rocha, W. Scheirer, T. Boult, and S. Goldenstein, \u201cVision of the", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 443 }, { "text": "Unseen: Current Trends and Challenges in Digital Image and Video Forensics,\u201d ACM Computing Surveys, vol. 43, no. 4, pp. 1\u201342, 2011. [9] S. Milani, M. Fontani, P. Bestagini, M. Barni, A. Piva, M. Tagliasacchi, and S. Tubaro, \u201cAn Overview on Video Forensics,\u201d APSIPA Transac- tions on Signal and Information Processing, vol. 1, pp. 1\u201318, 2012.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 444 }, { "text": "M. Nie\u00dfner, \u201cFaceForensics++: Learning to Detect Manipulated Facial Images,\u201d in Proc. IEEE/CVF International Conference on Computer Vision, 2019. [13] J. Galbally, S. Marcel, and J. Fierrez, \u201cBiometric Anti-Spoo\ufb01ng Meth- ods: A Survey in Face Recognition,\u201d IEEE Access, vol. 2, pp. 1530\u2013 1552, 2014.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 445 }, { "text": "of Digital Face Manipulation,\u201d in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020. [18] C. Canton, L. Davis, E. Delp, P. Flynn, S. McCloskey, L. Leal-Taixe, P. Natsev, and C. Bregler, \u201cApplications of Computer Vision and Pattern Recognition to Media Forensics,\u201d in IEEE/CVF Conference on", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 446 }, { "text": "Computer Vision and Pattern Recognition, 2019. [Online]. Available: https://sites.google.com/view/mediaforensics2019 [19] B. Biggio, P. Korshunov, T. Mensink, G. Patrini, D. Rao, and A. Sadhu, \u201cSynthetic Realities: Deep Learning for Detecting AudioVisual Fakes,\u201d in International Conference on Machine Learning, 2019. [Online].", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 447 }, { "text": "Available: https://sites.google.com/view/audiovisualfakes-icml2019/ [20] L. Verdoliva and P. Bestagini, \u201cMultimedia Forensics,\u201d in ACM Multimedia, 2019. [Online]. Available: https://acmmm.org/tutorials/ #tut3 [21] K. Raja, N. Damer, C. Chen, A. Dantcheva, A. Czajka, H. Han, and R. Ramachandra, \u201cWorkshop on Deepfakes and Presentation", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 448 }, { "text": "Attacks in Biometrics,\u201d in IEEE Winter Conference on Applications of Computer Vision, 2020. [Online]. Available: https://sites.google.com/ view/wacv2020-deeppab/ [22] M. Barni, S. Battiato, G. Boato, H. Farid, and N. Memon, \u201cMultiMedia Forensics in the Wild,\u201d in IEEE International Conference on Pattern Recognition, 2020. [Online]. Available: https://iplab.dmi.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 449 }, { "text": "unict.it/mmforwild/ [23] C. Bregler, M. Covell, and M. Slaney, \u201cVideo Rewrite: Driving Visual Speech with Audio,\u201d Computer Graphics, vol. 31, no. 2, pp. 353\u2013361, 1997. [24] D.P. Kingma and M. Welling, \u201cAuto-Encoding Variational Bayes,\u201d in Proc. International Conference on Learning Representations, 2013.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 450 }, { "text": "Conference on Computer Vision and Pattern Recognition, 2020. [27] I. Yerushalmy and H. Hel-Or, \u201cDigital Image Forgery Detection based on Lens and Sensor Aberration,\u201d International Journal of Computer Vision, vol. 92, no. 1, pp. 71\u201391, 2011. [28] A.C. Popescu and H. Farid, \u201cExposing Digital Forgeries in Color Filter", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 451 }, { "text": "Array Interpolated Images,\u201d IEEE Transactions on Signal Processing, vol. 53, no. 10, pp. 3948\u20133959, 2005. [29] H. Cao and A.C. Kot, \u201cAccurate Detection of Demosaicing Regular- ity for Digital Image Forensics,\u201d IEEE Transactions on Information Forensics and Security, vol. 4, no. 4, pp. 899\u2013910, 2009.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 452 }, { "text": "tection,\u201d IEEE Transactions on Information Forensics and Security, vol. 6, no. 2, pp. 396\u2013406, 2011. [32] I. Amerini, L. Ballan, R. Caldelli, A. Bimbo, and G. Serra, \u201cA SIFT- Based Forensic Method for Copy-Move Attack Detection and Trans- formation Recovery,\u201d IEEE Transactions on Information Forensics and", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 453 }, { "text": "Security, vol. 6, no. 3, pp. 1099\u20131110, 2011. [33] D. Cozzolino, G. Poggi, and L. Verdoliva, \u201cSplicebuster: A new blind image splicing detector,\u201d in Proc. IEEE International Workshop on Information Forensics and Security, 2015, pp. 1\u20136. [34] A. Gironi, M. Fontani, T. Bianchi, A. Piva, and M. Barni, \u201cA Video", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 454 }, { "text": "Forensic Technique for Detecting Frame Deletion and Insertion,\u201d in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing, 2014, pp. 6226\u20136230. [35] Y. Wu, X. Jiang, T. Sun, and W. Wang, \u201cExposing Video Inter- Frame Forgery based on Velocity Field Consistency,\u201d in Proc. IEEE", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 455 }, { "text": "19 [37] H. Allcott and M. Gentzkow, \u201cSocial Media and Fake News in the 2016 Election,\u201d Journal of Economic Perspectives, vol. 31, no. 2, pp. 211\u201336, 2017. [38] D. M. Lazer, M. A. Baum, Y. Benkler, A. J. Berinsky, K. M. Greenhill, F. Menczer, M. J. Metzger, B. Nyhan, G. Pennycook, D. Rothschild et al., \u201cThe Science of Fake News,\u201d Science, vol. 359, no. 6380, pp.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 456 }, { "text": "1094\u20131096, 2018. [39] I. M. J. Kietzmann, L.W. Lee and T. Kietzmann, \u201cDeepfakes: Trick or treat?\u201d Business Horizons, vol. 63, no. 2, pp. 135\u2013146, 2020. [40] L. Verdoliva, \u201cMedia Forensics and DeepFakes: an Overview,\u201d arXiv preprint arXiv:2001.06564, 2020. [41] T. Karras, S. Laine, and T. Aila, \u201cA Style-Based Generator Architecture", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 457 }, { "text": "for Generative Adversarial Networks,\u201d in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019. [42] E. Gonzalez-Sosa, J. Fierrez, R. Vera-Rodriguez, and F. Alonso- Fernandez, \u201cFacial Soft Biometrics for Recognition in the Wild: Recent Works, Annotation and COTS Evaluation,\u201d IEEE Transactions on", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 458 }, { "text": "Information Forensics and Security, vol. 13, no. 8, pp. 2001\u20132014, 2018. [43] Y. Choi, M. Choi, M. Kim, J. Ha, S. Kim, and J. Choo, \u201cStarGAN: Uni\ufb01ed Generative Adversarial Networks for Multi-Domain Image- to-Image Translation,\u201d in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 459 }, { "text": "\u201cFace2face: Real-Time Face Capture and Reenactment of RGB Videos,\u201d in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2016. [46] J. Thies, M. Zollh\u00a8ofer, and M. Nie\u00dfner, \u201cDeferred Neural Rendering: Image Synthesis using Neural Textures,\u201d ACM Transactions on Graph- ics, vol. 38, no. 66, pp. 1\u201312, 2019.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 460 }, { "text": "Leave Arti\ufb01cial Fingerprints?\u201d in Proc. IEEE Conference on Multime- dia Information Processing and Retrieval, 2019, pp. 506\u2013511. [50] M. Albright and S. McCloskey, \u201cSource Generator Attribution via Inversion?\u201d in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2019. [51] A. Jain, P. Majumdar, R. Singh, and M. Vatsa, \u201cDetecting GANs", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 461 }, { "text": "and Retouching based Digital Alterations via DAD-HCNN,\u201d in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2020. [52] Z. Liu, P. Luo, X. Wang, and X. Tang, \u201cDeep Learning Face Attributes in the Wild,\u201d in Proc. IEEE/CVF International Conference on Com- puter Vision, 2015.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 462 }, { "text": "2018. [55] S. McCloskey and M. Albright, \u201cDetecting GAN-Generated Imagery Using Color Cues,\u201d arXiv preprint arXiv:1812.08247, 2018. [56] R. Wang, L. Ma, F. Juefei-Xu, X. Xie, J. Wang, and Y. Liu, \u201cFakeSpot- ter: A Simple Baseline for Spotting AI-Synthesized Fake Faces,\u201d arXiv preprint arXiv:1909.06122, 2019.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 463 }, { "text": "Fake Images Using Co-Occurrence Matrices,\u201d Electronic Imaging, no. 5, pp. 1\u20137, 2019. [59] N. Yu, L. Davis, and M. Fritz, \u201cAttributing Fake Images to GANs: Analyzing Fingerprints in Generated Images,\u201d in Proc. IEEE/CVF International Conference on Computer Vision, 2019. [60] F. Marra, C. Saltori, G. Boato, and L. Verdoliva, \u201cIncremental Learning", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 464 }, { "text": "for the Detection and Classi\ufb01cation of GAN-Generated Images,\u201d in Proc. IEEE International Workshop on Information Forensics and Security, 2019. [61] N. Hulzebosch, Sarah Ibrahimi and Marcel Worring, \u201cDetecting CNN- Generated Facial Images in Real-World Scenarios,\u201d in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops,", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 465 }, { "text": "2020. [62] H. Guan, M. Kozak, E. Robertson, Y. Lee, A. Yates, A. Delgado, D. Zhou, T. Kheyrkhah, J. Smith, and J. Fiscus, \u201cMFC Datasets: Large- Scale Benchmark Datasets for Media Forensic Challenge Evaluation,\u201d in Proc. IEEE Winter Applications of Computer Vision Workshops, 2019. [63] O. Parkhi, A. Vedaldi, and A. Zisserman, \u201cDeep Face Recognition,\u201d in", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 466 }, { "text": "Proc. British Machine Vision Conference, 2015. [64] B. Amos, B. Ludwiczuk, and M. Satyanarayanan, \u201cOpenFace: A General-Purpose Face Recognition Library with Mobile Applications,\u201d in CMU School of Computer Science, 2016. [65] F. Schroff, D. Kalenichenko, and J. Philbin, \u201cFaceNet: A Uni\ufb01ed Embedding for Face Recognition and Clustering,\u201d in Proc. IEEE/CVF", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 467 }, { "text": "Conference on Computer Vision and Pattern Recognition, 2015. [66] Y. Shen, J. Gu, X. Tang, and B. Zhou, \u201cInterpreting the Latent Space of GANs for Semantic Face Editing,\u201d in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020. [67] T. K. Moon, \u201cThe Expectation-Maximization Algorithm,\u201d IEEE Signal", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 468 }, { "text": "Processing Magazine, vol. 13, no. 6, pp. 47\u201360, 1996. [68] Z. He, W. Zuo, M. Kan, S. Shan, and X. Chen, \u201cAttGAN: Facial At- tribute Editing by Only Changing What You Want,\u201d IEEE Transactions on Image Processing, vol. 28, no. 11, pp. 5464\u20135478, 2019. [69] W. Cho, S. Choi, D. K. Park, I. Shin, and J. Choo, \u201cImage-to-Image", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 469 }, { "text": "Translation via Group-Wise Deep Whitening-and-Coloring Transfor- mation,\u201d in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019. [70] T. Karras, S. Laine, M. Aittala, J. Hellsten, J. Lehtinen, and T. Aila, \u201cAnalyzing and Improving the Image Quality of StyleGAN,\u201d in Proc. IEEE/CVF Conference on Computer Vision and Patter Recognition,", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 470 }, { "text": "2020. [71] J. Zhu, T. Park, P. Isola, and A. Efros, \u201cUnpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks,\u201d in Proc. IEEE/CVF International Conference on Computer Vision, 2017. [72] T. Miyato, T. Kataoka, M. Koyama, and Y. Yoshida, \u201cSpectral Normal- ization for Generative Adversarial Networks,\u201d in Proc. International", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 471 }, { "text": "Conference on Learning Representations, 2018. [73] M. Bellemare, I. Danihelka, W. Dabney, S. Mohamed, B. Lak- shminarayanan, S. Hoyer, and R. Munos, \u201cThe Cramer Distance as a Solution to Biased Wasserstein Gradients,\u201d arXiv preprint arXiv:1705.10743, 2017. [74] M. Binkowski, D. Sutherland, M. Arbel, and A. Gretton, \u201cDemysti-", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 472 }, { "text": "fying MMD GANs,\u201d in Proc. International Conference on Learning Representations, 2018. [75] S. Rebuf\ufb01, A. Kolesnikov, G. Sperl, and C. Lampert, \u201ciCaRL: Incre- mental Classi\ufb01er and Representation Learning,\u201d in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2017. [76] D. Kingma and P. Dhariwal, \u201cGlow: Generative Flow with Invertible", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 473 }, { "text": "1x1 Convolutions,\u201d in Proc. Advances in Neural Information Process- ing Systems, 2018. [77] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. Gomez, L. Kaiser, and I. Polosukhin, \u201cAttention is All You Need,\u201d in Proc. Advances in Neural Information Processing Systems, 2017, pp. 5998\u2013 6008.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 474 }, { "text": "Computer Vision and Pattern Recognition, 2018. [82] Y. Li, M. Chang, and S. Lyu, \u201cIn Ictu Oculi: Exposing AI Generated Fake Face Videos by Detecting Eye Blinking,\u201d in Proc. IEEE Interna- tional Workshop on Information Forensics and Security, 2018. [83] B. Dolhansky, R. Howes, B. P\ufb02aum, N. Baram, and C. Ferrer,", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 475 }, { "text": "20 [85] K. Zhang, Z. Zhang, Z. Li, and Y. Qiao, \u201cJoint Face Detection and Alignment Using Multitask Cascaded Convolutional Networks,\u201d IEEE Signal Processing Letters, vol. 23, no. 10, pp. 1499\u20131503, 2016. [86] Google AI, \u201cContributing Data to Deepfake Detection Research,\u201d 2019. [Online]. Available: https://ai.googleblog.com/2019/", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 476 }, { "text": "09/contributing-data-to-deepfake-detection.html [87] A. R\u00a8ossler, D. Cozzolino, L. Verdoliva, C. Riess, J. Thies, and M. Nie\u00dfner, \u201cFaceForensics: A Large-Scale Video Dataset for Forgery Detection in Human Faces,\u201d arXiv preprint arXiv:1803.09179, 2018. [88] P. P\u00b4erez, M. Gangnet, and A. Blake, \u201cPoisson Image Editing,\u201d ACM", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 477 }, { "text": "Transactions on Graphics, vol. 22, no. 3, pp. 313\u2013318, 2003. [89] F. Matern, C. Riess, and M. Stamminger, \u201cExploiting Visual Artifacts to Expose DeepFakes and Face Manipulations,\u201d in Proc. IEEE Winter Applications of Computer Vision Workshops, 2019. [90] X. Yang, Y. Li, and S. Lyu, \u201cExposing Deep Fakes Using Inconsistent", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 478 }, { "text": "Head Poses,\u201d in Proc. International Conference on Acoustics, Speech and Signal Processing, 2019. [91] S. Agarwal and H. Farid, \u201cProtecting World Leaders Against Deep Fakes,\u201d in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2019. [92] T. Jung, S. Kim, and K. Kim, \u201cDeepVision: Deepfakes Detection Using", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 479 }, { "text": "Human Eye Blinking Pattern,\u201d IEEE Access, vol. 8, pp. 83 144\u201383 154, 2020. [93] Y. Li and S. Lyu, \u201cExposing DeepFake Videos By Detecting Face Warping Artifacts,\u201d in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2019. [94] D. Afchar, V. Nozick, J. Yamagishi, and I. Echizen, \u201cMesoNet: a", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 480 }, { "text": "Compact Facial Video Forgery Detection Network,\u201d in Proc. IEEE International Workshop on Information Forensics and Security, 2018. [95] P. Zhou, X. Han, V. Morariu, and L. Davis, \u201cTwo-Stream Neural Net- works for Tampered Face Detection,\u201d in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2017.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 481 }, { "text": "2019. [98] E. Sabir, J. Cheng, A. Jaiswal, W. AbdAlmageed, I. Masi, and P. Natarajan, \u201cRecurrent Convolutional Strategies for Face Manipula- tion Detection in Videos,\u201d in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2019. [99] D. G\u00a8uera and E. Delp, \u201cDeepfake Video Detection Using Recurrent", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 482 }, { "text": "Neural Networks,\u201d in Proc. International Conference on Advanced Video and Signal Based Surveillance, 2018. [100] R. Tolosana, S. Romero-Tapiador, J. Fierrez and R. Vera-Rodriguez, \u201cDeepFakes Evolution: Analysis of Facial Regions and Fake Detection Performance,\u201d arXiv preprint arXiv:2004.07532, 2020.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 483 }, { "text": "pp. 710\u2013724, 2014. [103] I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning, 2016. [104] D. King, \u201cDLib-ML: A Machine Learning Toolkit,\u201d Journal of Machine Learning Research, vol. 10, pp. 1755\u20131758, 2009. [105] T. Baltrusaitis, A. Zadeh, Y. Lim, and L. Morency, \u201cOpenFace 2.0: Facial Behavior Analysis Toolkit,\u201d in Proc. International Conference", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 484 }, { "text": "on Automatic Face & Gesture Recognition, 2018. [106] R. Daza, A. Morales, J. Fierrez, and R. Tolosana, \u201cmEBAL: A Multimodal Database for Eye Blink Detection and Attention Level Estimation,\u201d arXiv preprint arXiv:2006.05327, 2020. [107] R. Ranjan, V. M. Patel, and R. Chellappa, \u201cHyperface: A deep Multi-", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 485 }, { "text": "Task Learning Framework for Face Detection, Landmark Localization, Pose Estimation, and Gender Recognition,\u201d IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 41, no. 1, pp. 121\u2013 135, 2017. [108] T. Soukupov\u00b4a and J. Cech, \u201cEye Blink Detection Using Facial Land- marks,\u201d in Proc. Computer Vision Winter Workshop, 2016.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 486 }, { "text": "Convolutions,\u201d in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2015. [113] D. Cozzolino, G. Poggi, and L. Verdoliva, \u201cRecasting Residual-Based Local Descriptors as Convolutional Neural Networks: an Application to Image Forgery Detection,\u201d in Proc. ACM Workshop on Information", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 487 }, { "text": "Hiding and Multimedia Security, 2017. [114] B. Bayar and M. Stamm, \u201cA Deep Learning Approach to Universal Image Manipulation Detection Using a New Convolutional Layer,\u201d in Proc. ACM Workshop on Information Hiding and Multimedia Security, 2016. [115] N. Rahmouni, V. Nozick, J. Yamagishi, and I. Echizen, \u201cDistinguishing", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 488 }, { "text": "Computer Graphics from Natural Images Using Convolution Neural Networks,\u201d in Proc. IEEE Workshop on Information Forensics and Security, 2017. [116] F. Chollet, \u201cXception: Deep Learning with Depthwise Separable Con- volutions,\u201d in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2017.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 489 }, { "text": "2011, pp. 44\u201351. [119] S. Sabour, N. Frosst and G.E. Hinton, \u201cDynamic Routing Between Cap- sules,\u201d in Proc. Advances in Neural Information Processing Systems, 2017, pp. 3856\u20133866. [120] G.E. Hinton, S. Sabour and N. Frosst, \u201cMatrix Capsules with EM rout- ing,\u201d in Proc. International Conference on Learning Representations", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 490 }, { "text": "Workshop, 2018. [121] Y. Wang and A. Dantcheva, \u201cA Video is Worth More than 1000 Lies. Comparing 3DCNN Approaches for Detecting Deepfakes,\u201d in Proc. IEEE International Conference on Automatic Face and Gesture Recognition, 2020. [122] J. Carreira and A. Zisserman, \u201cQuo Vadis, Action Recognition? A New", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 491 }, { "text": "Model and the Kinetics Dataset,\u201d in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2017. [123] K. Hara, H. Kataoka, and Y. Satoh, \u201cCan Spatiotemporal 3D CNNs Retrace the History of 2D CNNs and Imagenet?\u201d in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 492 }, { "text": "Graphics, vol. 36, no. 4, pp. 1\u201313, 2017. [126] P. Isola, J. Zhu, T. Zhou, and A. Efros, \u201cImage-to-Image Translation with Conditional Adversarial Networks,\u201d in Proc. IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, 2017. [127] J. Zhu, R. Zhang, D. Pathak, T. Darrell, A. Efros, O. Wang, and", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 493 }, { "text": "E. Shechtman, \u201cToward Multimodal Image-to-Image Translation,\u201d in Proc. Advances in Neural Information Processing Systems, 2017. [128] T. Kim, M. Cha, H. Kim, J. Lee, and J. Kim, \u201cLearning to Discover Cross-Domain Relations with Generative Adversarial Networks,\u201d in Proc. International Conference on Machine Learning, 2017.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 494 }, { "text": "Information Processing Systems Workshops, 2016. [131] M. Li, W. Zuo, and D. Zhang, \u201cDeep Identity-Aware Transfer of Facial Attributes,\u201d arXiv preprint arXiv:1610.05586, 2016. [132] W. Shen and R. Liu, \u201cLearning Residual Images for Face Attribute Manipulation,\u201d in Proc. IEEE/CVF Conference on Computer Vision", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 495 }, { "text": "21 [134] T. Xiao, J. Hong, and J. Ma, \u201cELEGANT: Exchanging Latent Encod- ings with GAN for Transferring Multiple Face Attributes,\u201d in Proc. European Conference on Computer Vision, 2018. [135] M. Mirza and S. S. Osindero, \u201cConditional Generative Adversarial Nets,\u201d arXiv preprint arXiv:1411.1784, 2014.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 496 }, { "text": "Biometrics Theory, Applications and Systems, 2018. [138] S. Tariq, S. Lee, H. Kim, Y. Shin, and S. Woo, \u201cDetecting Both Machine and Human Created Fake Face Images in the Wild,\u201d in Proc. International Workshop on Multimedia Privacy and Security, 2018, pp. 81\u201387. [139] S. Wang, O. Wang, A. Owens, R. Zhang, and A. Efros, \u201cDetecting", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 497 }, { "text": "Photoshopped Faces by Scripting Photoshop,\u201d in Proc. IEEE/CVF International Conference on Computer Vision, 2019. [140] X. Zhang, S. Karaman, and S. Chang, \u201cDetecting and Simulating Artifacts in GAN Fake Images,\u201d in Proc. IEEE International Workshop on Information Forensics and Security, 2019. [141] C. Rathgeb, A. Botaljov, F. Stockhardt, S. Isadskiy, L. Debiasi, A. Uhl,", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 498 }, { "text": "and C. Busch, \u201cPRNU-based Detection of Facial Retouching,\u201d IET Biometrics, 2020. [142] J. Kim, J. Choi, J. Yi, and M. Turk, \u201cEffective Representation Using ICA for Face Recognition Robust to Local Distortion and Partial Occlusion,\u201d IEEE Transactions on Pattern Analysis and Machine Intelligence, no. 12, pp. 1977\u20131981, 2005.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 499 }, { "text": "tional Conference and Workshops on Automatic Face and Gesture Recognition, 2015. [145] P. Majumdar, A. Agarwal, R. Singh, and M. Vatsa, \u201cEvading Face Recognition via Partial Tampering of Faces,\u201d in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2019. [146] C. Rathgeb, A. Dantcheva, and C. Busch, \u201cImpact and Detection of", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 500 }, { "text": "Facial Beauti\ufb01cation in Face Recognition: An Overview,\u201d IEEE Access, vol. 7, pp. 152 667\u2013152 678, 2019. [147] C. Rathgeb, C. I. Satnoianu, N. E. Haryanto, K. Bernardo, and C. Busch, \u201cDifferential Detection of Facial Retouching: A Multi- Biometric Approach,\u201d IEEE Access, vol. 8, pp. 106 373\u2013106 385, 2020.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 501 }, { "text": "face-aware-liquify.html [150] D. Sun, X. Yang, M.Y. Liu and J. Kautz, \u201cPWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume,\u201d in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018. [151] H. Averbuch-Elor, D. Cohen-Or, J. Kopf and M.F. Cohen, \u201cBringing Portraits to Life,\u201d ACM Transactions on Graphics, vol. 36, no. 6, p.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 502 }, { "text": "196, 2017. [152] E. Zakharov, A. Shysheya, E. Burkov, and V. Lempitsky, \u201cFew-Shot Adversarial Learning of Realistic Neural Talking Head Models,\u201d in Proc. IEEE/CVF International Conference on Computer Vision, 2019. [153] D. Zhu, S. Liu, W. Jiang, C. Gao, T. Wu, and G. Guo, \u201cUGAN: Untraceable GAN for Multi-Domain Face Translation,\u201d arXiv preprint", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 503 }, { "text": "arXiv:1907.11418, 2019. [154] I. Amerini, L. Galteri, R. Caldelli, and A. Bimbo, \u201cDeepfake Video Detection through Optical Flow based CNN,\u201d in Proc. IEEE/CVF International Conference on Computer Vision, 2019. [155] G. Wolberg, \u201cImage Morphing: a Survey,\u201d The Visual Computer, vol. 14, no. 8-9, pp. 360\u2013372, 1998.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 504 }, { "text": "International Workshop on Biometrics and Forensics, 2017. [158] P. Korshunov and S. Marcel, \u201cVulnerability of Face Recognition to Deep Morphing,\u201d arXiv preprint arXiv:1910.01933, 2019. [159] K. Raja, M. Ferrara, A. Franco, L. Spreeuwers, I. Batskos, F. de Wit, M. Gomez-Barrero, U. Scherhag, D. Fischer, S. Venkatesh, J. M.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 505 }, { "text": "Singh, G. Li, L. Bergeron, S. Isadskiy, R. Ramachandra, C. Rathgeb, D. Frings, U. Seidel, F. Knopjes, R. Veldhuis, D. Maltoni, and C. Busch, \u201cMorphing Attack Detection - Database, Evaluation Platform and Benchmarking,\u201d arXiv preprint arXiv:2006.06458, 2020. [160] C. Kraetzer, A. Makrushin, T. Neubert,", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 506 }, { "text": "M. Hildebrandt, and J. Dittmann, \u201cModeling Attacks on Photo-ID Documents and Applying Media Forensics for the Detection of Facial Morphing,\u201d in Proc. ACM Workshop on Information Hiding and Multimedia Security, 2017, pp. 21\u201332. [161] L.B. Zhang, F. Peng and M. Long, \u201cFace Morphing Detection Using Fourier Spectrum of Sensor Pattern Noise,\u201d in Proc. International", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 507 }, { "text": "Conference on Multimedia and Expo, 2018. [162] U. Scherhag, D. Budhrani, M. Gomez-Barrero, and C. Busch, \u201cDetect- ing Morphed Face Images Using Facial Landmarks,\u201d in Proc. IEEE International Conference on Image and Signal Processing, 2018. [163] N. Damer, V. Bolle, Y. Wainakh, F. Boutros, P. Terh\u00a8orst, A. Braun,", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 508 }, { "text": "and A. Kuijper, \u201cDetecting Face Morphing Attacks by Analyzing the Directed Distances of Facial Landmarks Shifts,\u201d in German Conference on Pattern Recognition, 2018. [164] M. Ferrara, A. Franco, and D. Maltoni, \u201cFace Morphing Detection in the Presence of Printing/Scanning and Heterogeneous Image Sources,\u201d", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 509 }, { "text": "arXiv preprint arXiv:1901.08811, 2019. [165] U. Scherhag, C. Rathgeb, J. Merkle, and C. Busch, \u201cDeep Face Representations for Differential Morphing Attack Detection,\u201d arXiv preprint arXiv:2001.01202, 2020. [166] M. Ferrara, A. Franco, and D. Maltoni, \u201cFace demorphing,\u201d IEEE Transactions on Information Forensics and Security, vol. 13, no. 4,", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 510 }, { "text": "pp. 1008\u20131017, 2017. [167] F. Peng, L.B. Zhang and M. Long, Min, \u201cFD-GAN: Face De-Morphing Generative Adversarial Network for Restoring Accomplices Facial Image,\u201d IEEE Access, vol. 7, pp. 75 122\u201375 131, 2019. [168] R. Gross, L. Sweeney, F. De la Torre, and S. Baker, \u201cModel-Based Face De-Identi\ufb01cation,\u201d in Proc. IEEE/CVF Conference on Computer", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 511 }, { "text": "Vision and Pattern Recognition Workshop, 2006. [169] R. Gross, L. Sweeney, J. Cohn, F. De la Torre, and S. Baker, \u201cFace De- Identi\ufb01cation,\u201d in Protecting Privacy in Video Surveillance. Springer, 2009, pp. 129\u2013146. [170] B. Meden, R. C. Mall\u0131, S. Fabijan, H. K. Ekenel, V. \u02c7Struc, and P. Peer, \u201cFace Deidenti\ufb01cation with Generative Deep Neural Networks,\u201d IET", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 512 }, { "text": "Signal Processing, vol. 11, no. 9, pp. 1046\u20131054, 2017. [171] K. Brkic, I. Sikiric, T. Hrkac, and Z. Kalafatic, \u201cI Know That Person: Generative Full Body and Face De-identi\ufb01cation of People in Images,\u201d in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2017. [172] Q. Sun, L. Ma, S.O. Joon, L.V. Gool, B. Schiele and M. Fritz, \u201cNatural", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 513 }, { "text": "and Effective Obfuscation by Head Inpainting,\u201d in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018. [173] B. Meden, V. Emer\u02c7si\u02c7c, V. \u02c7Struc, and P. Peer, \u201ck-Same-Net: k- Anonymity with Generative Deep Neural Networks for Face Deiden- ti\ufb01cation,\u201d Entropy, vol. 20, no. 1, p. 60, 2018.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 514 }, { "text": "and Mixed Precision Training,\u201d in Proc. International Conference on Advanced Video and Signal Based Surveillance, 2019. [176] V. Mirjalili, S. Raschka, and A. Ross, \u201cFlowSAN: Privacy-Enhancing Semi-Adversarial Networks to Confound Arbitrary Face-based Gender Classi\ufb01ers,\u201d IEEE Access, vol. 7, pp. 99 735\u201399 745, 2019.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 515 }, { "text": "Explicit Removal of Biases and Variation from Deep Neural Network Embeddings,\u201d in Proc. European Conference on Computer Vision, 2018. [180] A. Morales, J. Fierrez, and R. Vera-Rodriguez, \u201cSensitiveNets: Learn- ing Agnostic Representations with Application to Face Recognition,\u201d arXiv preprint arXiv:1902.00334, 2019.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 516 }, { "text": "22 [181] S. Gong, X. Liu, and A. Jain, \u201cDebFace: De-biasing Face Recognition,\u201d arXiv preprint arXiv:1911.08080, 2019. [182] S. Agarwal, H. Farid, O. Fried, and M. Agrawala, \u201cDetecting Deep- Fake Videos from Phoneme-Viseme Mismatches,\u201d in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops,", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 517 }, { "text": "2020. [183] Y. Song, J. Zhu, D. Li, A. Wang, and H. Qi, \u201cTalking Face Generation by Conditional Recurrent Adversarial Network,\u201d in Proc. International Joint Conference on Arti\ufb01cial Intelligence, 2019. [184] L. Song, W. Wu, C. Qian, R. He, and C. Loy, \u201cEverybody\u2019s Talkin\u2019: Let Me Talk as You Want,\u201d arXiv preprint arXiv:2001.05201, 2020.", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 518 }, { "text": "\u201cText-Based Editing of Talking-Head Video,\u201d ACM Transactions on Graphics, vol. 38, no. 4, pp. 1\u201314. [187] H. Khalid and S. S. Woo, \u201cOC-FakeDect: Classifying Deepfakes Using One-class Variational Autoencoder,\u201d in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2020. [188] S. Fernandes, S. Raj, R. Ewetz, J. S. Pannu, S. K. Jha, E. Ortiz,", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 519 }, { "text": "I. Vintila, and M. Salter, \u201cDetecting Deepfake Videos using Attribution- Based Con\ufb01dence Metric,\u201d in Proc. IEEE/CVF Conference on Com- puter Vision and Pattern Recognition Workshops, 2020. [189] J. Fierrez, A. Morales, R. Vera-Rodriguez, and D. Camacho, \u201cMultiple Classi\ufb01ers in Biometrics. Part 1: Fundamentals and Review,\u201d Informa-", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 520 }, { "text": "tion Fusion, vol. 44, pp. 57\u201364, 2018. [190] \u2014\u2014, \u201cMultiple Classi\ufb01ers in Biometrics. Part 2: Trends and Chal- lenges,\u201d Information Fusion, vol. 44, pp. 103\u2013112, 2018. [191] R. S. M. Singh and A. Ross, \u201cA Comprehensive Overview of Biometric Fusion,\u201d Information Fusion, vol. 52, pp. 187\u2013205, 2019. [192] Q. Yang, X. Zhu, J. K. Fwu, Y. Ye, G. You, and Y. Zhu, \u201cPipeNet:", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 521 }, { "text": "Selective Modal Pipeline of Fusion Network for Multi-Modal Face Anti-Spoo\ufb01ng,\u201d in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2020. [193] Z. Yu, Y. Qin, X. Li, Z. Wang, C. Zhao, Z. Lei, and G. Zhao, \u201cMulti- Modal Face Anti-Spoo\ufb01ng Based on Central Difference Networks,\u201d", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 522 }, { "text": "in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2020. [194] D. M. Montserrat, H. Hao, S. K. Yarlagadda, S. Baireddy, R. Shao, J. Horvath, E. Bartusiak, J. Yang, D. G\u00a8uera, F. Zhu, and E. J. Delp, \u201cDeepfakes Detection with Automatic Face Weighting,\u201d in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 523 }, { "text": "Workshops, 2020. [195] T. Agrawal, R. Gupta, and S. Narayanan, \u201cMultimodal Detection of Fake Social Media Use through a Fusion of classi\ufb01cation and Pairwise Ranking Systems,\u201d in Proc. European Signal Processing Conference, 2017, pp. 1045\u20131049. [196] K. Shu, A. Sliva, S. Wang, J. Tang, and H. Liu, \u201cFake News Detection", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 524 }, { "text": "on Social Media: A Data Mining Perspective,\u201d ACM SIGKDD Explo- rations Newsletter, vol. 19, no. 1, pp. 22\u201336, 2017. [197] K. Shu, D. Mahudeswaran, and H. Liu, \u201cFakeNewsTracker: a Tool for Fake News Collection, Detection, and Visualization,\u201d Computational and Mathematical Organization Theory, vol. 25, no. 1, pp. 60\u201371,", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 525 }, { "text": "2019. [198] A. Morales, A. Acien, J. Fierrez, J. V. Monaco, R. Tolosana, R. Vera- Rodriguez, and J. Ortega-Garcia, \u201cKeystroke Biometrics in Response to Fake News Propagation in a Global Pandemic,\u201d in Proc. IEEE Computer Software and Applications Conference Workshops, 2020. [199] E. Tursman, M. George, S. Kamara, and J. Tompkin, \u201cTowards", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 526 }, { "text": "Untrusted Social Video Veri\ufb01cation to Combat Deepfakes via Face Geometry Consistency,\u201d in Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2020. [200] N. Carlini and H. Farid, \u201cEvading Deepfake-Image Detectors with White- and Black-Box Attacks,\u201d in Proc. IEEE/CVF Conference on", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 527 }, { "text": "Computer Vision and Pattern Recognition Workshops, 2020. Ruben Tolosana received the M.Sc. degree in Telecommunication Engineering, and his Ph.D. de- gree in Computer and Telecommunication Engineer- ing, from Universidad Autonoma de Madrid, in 2014 and 2019, respectively. In April 2014, he joined the", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 528 }, { "text": "Biometrics and Data Pattern Analytics - BiDA Lab at the Universidad Autonoma de Madrid, where he is currently collaborating as a PostDoctoral researcher. Since then, Ruben has been granted with several awards such as the FPU research fellowship from Spanish MECD (2015), and the European Biomet- rics Industry Award (2018). His research interests are mainly focused on signal", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 529 }, { "text": "and image processing, pattern recognition, and machine learning, particularly in the areas of face manipulation, human-computer interaction and biometrics. He is author of several publications and also collaborates as a reviewer in many different high-impact conferences (e.g., ICDAR, IJCB, ICB, BTAS,", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 530 }, { "text": "EUSIPCO, etc.) and journals (e.g., IEEE TPAMI, TCYB, TIFS, TIP, ACM CSUR, etc.). Finally, he has participated in several National and European projects focused on the deployment of biometric security through the world. Ruben Vera-Rodriguez received the M.Sc. degree in telecommunications engineering from Universi-", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 531 }, { "text": "dad de Sevilla, Spain, in 2006, and the Ph.D. de- gree in electrical and electronic engineering from Swansea University, U.K., in 2010. Since 2010, he has been af\ufb01liated with the Biometric Recognition Group, Universidad Autonoma de Madrid, Spain, where he is currently an Associate Professor since 2018. His research interests include signal and image", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 532 }, { "text": "processing, pattern recognition, and biometrics, with emphasis on signature, face, gait veri\ufb01cation and forensic applications of biometrics. He is actively involved in several National and European projects focused on biometrics. Ruben has been Program Chair for the IEEE 51st International Carnahan Conference on Security and", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 533 }, { "text": "Technology (ICCST) in 2017; and the 23rd Iberoamerican Congress on Pattern Recognition (CIARP 2018) in 2018. Julian Fierrez received the M.Sc. and Ph.D. de- grees in telecommunications engineering from the Universidad Politecnica de Madrid, Spain, in 2001 and 2006, respectively. Since 2002, he has been", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 534 }, { "text": "with the Biometric Recognition Group, Universidad Politecnica de Madrid. Since 2004, he has been with the Universidad Autonoma de Madrid, where he is currently an Associate Professor. From 2007 to 2009, he was a Visiting Researcher with Michigan State University, USA, under a Marie Curie Fel- lowship. His research interests include signal and", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 535 }, { "text": "image processing, pattern recognition, and biometrics, with an emphasis on multibiometrics, biometric evaluation, system security, forensics, and mobile applications of biometrics. He has been actively involved in multiple EU projects focused on biometrics (e.g., TABULA RASA and BEAT), and has attracted notable impact for his research. He was a recipient of a number of", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 536 }, { "text": "distinctions, including the EAB European Biometric Industry Award 2006, the EURASIP Best Ph.D. Award 2012, the Miguel Catalan Award to the Best Re- searcher under 40 in the Community of Madrid in the general area of science and technology, and the 2017 IAPR Young Biometrics Investigator Award. He is an Associate Editor of the IEEE TRANSACTIONS ON INFORMATION", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 537 }, { "text": "23 Aythami Morales received the M.Sc. degree in telecommunication engineering from the Universi- dad de Las Palmas de Gran Canaria in 2006 and the Ph.D. degree from La Universidad de Las Pal- mas de Gran Canaria in 2011. Since 2017, he is Associate Professor with the Universidad Autonoma de Madrid. He has conducted research stays at", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 538 }, { "text": "the Biometric Research Laboratory, Michigan State University, the Biometric Research Center, Hong Kong Polytechnic University, the Biometric System Laboratory, University of Bologna, and the Schepens Eye Research Institute. He has authored over 70 scienti\ufb01c articles published in international journals and conferences. He has participated in national", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 539 }, { "text": "and EU projects in collaboration with other universities and private entities, such as UAM, UPM, EUPMt, Indra, Union Fenosa, Soluziona, or Accenture. His research interests are focused on pattern recognition, computer vision, machine learning, and biometrics signal processing. He has received awards", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 540 }, { "text": "from the ULPGC, La Caja de Canarias, SPEGC, and COIT. Javier Ortega-Garcia received the M.Sc. degree in electrical engineering and the Ph.D. degree (cum laude) in electrical engineering from Universidad Politecnica de Madrid, Spain, in 1989 and 1996, respectively. He is currently a Full Professor at the", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 541 }, { "text": "Signal Processing Chair in Universidad Autonoma de Madrid - Spain, where he holds courses on biometric recognition and digital signal processing. He is a founder and Director of the BiDA-Lab, Biometrics and Data Pattern Analytics Group. He has authored over 300 international contributions, including book chapters, refereed journal, and conference papers. His research", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 542 }, { "text": "interests are focused on biometric pattern recognition (on-line signature veri\ufb01cation, speaker recognition, human-device interaction) for security, e- health and user pro\ufb01ling applications. He chaired Odyssey-04, The Speaker Recognition Workshop, ICB-2013, the 6th IAPR International Conference on Biometrics, and ICCST2017, the 51st IEEE International Carnahan Confer-", "source": "DeepFakes and Beyond Survey", "year": 2020, "url": "https://arxiv.org/abs/2001.00179", "id": 543 }, { "text": "Xception: Deep Learning with Depthwise Separable Convolutions Franc\u00b8ois Chollet Google, Inc. fchollet@google.com Abstract We present an interpretation of Inception modules in con- volutional neural networks as being an intermediate step in-between regular convolution and the depthwise separable convolution operation (a depthwise convolution followed by", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 544 }, { "text": "a pointwise convolution). In this light, a depthwise separable convolution can be understood as an Inception module with a maximally large number of towers. This observation leads us to propose a novel deep convolutional neural network architecture inspired by Inception, where Inception modules have been replaced with depthwise separable convolutions.", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 545 }, { "text": "We show that this architecture, dubbed Xception, slightly outperforms Inception V3 on the ImageNet dataset (which Inception V3 was designed for), and signi\ufb01cantly outper- forms Inception V3 on a larger image classi\ufb01cation dataset comprising 350 million images and 17,000 classes. Since the Xception architecture has the same number of param-", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 546 }, { "text": "eters as Inception V3, the performance gains are not due to increased capacity but rather to a more ef\ufb01cient use of model parameters. 1. Introduction Convolutional neural networks have emerged as the mas- ter algorithm in computer vision in recent years, and de- veloping recipes for designing them has been a subject of", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 547 }, { "text": "considerable attention. The history of convolutional neural network design started with LeNet-style models [10], which were simple stacks of convolutions for feature extraction and max-pooling operations for spatial sub-sampling. In 2012, these ideas were re\ufb01ned into the AlexNet architec- ture [9], where convolution operations were being repeated", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 548 }, { "text": "multiple times in-between max-pooling operations, allowing the network to learn richer features at every spatial scale. What followed was a trend to make this style of network increasingly deeper, mostly driven by the yearly ILSVRC competition; \ufb01rst with Zeiler and Fergus in 2013 [25] and then with the VGG architecture in 2014 [18].", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 549 }, { "text": "At this point a new style of network emerged, the Incep- tion architecture, introduced by Szegedy et al. in 2014 [20] as GoogLeNet (Inception V1), later re\ufb01ned as Inception V2 [7], Inception V3 [21], and most recently Inception-ResNet [19]. Inception itself was inspired by the earlier Network- In-Network architecture [11]. Since its \ufb01rst introduction,", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 550 }, { "text": "Inception has been one of the best performing family of models on the ImageNet dataset [14], as well as internal datasets in use at Google, in particular JFT [5]. The fundamental building block of Inception-style mod- els is the Inception module, of which several different ver- sions exist. In \ufb01gure 1 we show the canonical form of an", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 551 }, { "text": "Inception module, as found in the Inception V3 architec- ture. An Inception model can be understood as a stack of such modules. This is a departure from earlier VGG-style networks which were stacks of simple convolution layers. While Inception modules are conceptually similar to con- volutions (they are convolutional feature extractors), they", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 552 }, { "text": "empirically appear to be capable of learning richer repre- sentations with less parameters. How do they work, and how do they differ from regular convolutions? What design strategies come after Inception? 1.1. The Inception hypothesis A convolution layer attempts to learn \ufb01lters in a 3D space, with 2 spatial dimensions (width and height) and a chan-", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 553 }, { "text": "nel dimension; thus a single convolution kernel is tasked with simultaneously mapping cross-channel correlations and spatial correlations. This idea behind the Inception module is to make this process easier and more ef\ufb01cient by explicitly factoring it into a series of operations that would independently look at", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 554 }, { "text": "cross-channel correlations and at spatial correlations. More precisely, the typical Inception module \ufb01rst looks at cross- channel correlations via a set of 1x1 convolutions, mapping the input data into 3 or 4 separate spaces that are smaller than the original input space, and then maps all correlations in", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 555 }, { "text": "these smaller 3D spaces, via regular 3x3 or 5x5 convolutions. This is illustrated in \ufb01gure 1. In effect, the fundamental hy- pothesis behind Inception is that cross-channel correlations and spatial correlations are suf\ufb01ciently decoupled that it is preferable not to map them jointly 1. 1A variant of the process is to independently look at width-wise corre-", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 556 }, { "text": "Consider a simpli\ufb01ed version of an Inception module that only uses one size of convolution (e.g. 3x3) and does not include an average pooling tower (\ufb01gure 2). This Incep- tion module can be reformulated as a large 1x1 convolution followed by spatial convolutions that would operate on non- overlapping segments of the output channels (\ufb01gure 3). This", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 557 }, { "text": "observation naturally raises the question: what is the ef- fect of the number of segments in the partition (and their size)? Would it be reasonable to make a much stronger hypothesis than the Inception hypothesis, and assume that cross-channel correlations and spatial correlations can be mapped completely separately?", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 558 }, { "text": "map the spatial correlations of every output channel. This is shown in \ufb01gure 4. We remark that this extreme form of an Inception module is almost identical to a depthwise sepa- rable convolution, an operation that has been used in neural lations and height-wise correlations. This is implemented by some of the", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 559 }, { "text": "modules found in Inception V3, which alternate 7x1 and 1x7 convolutions. The use of such spatially separable convolutions has a long history in im- age processing and has been used in some convolutional neural network implementations since at least 2012 (possibly earlier). Figure 3. A strictly equivalent reformulation of the simpli\ufb01ed In-", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 560 }, { "text": "ception module. Figure 4. An \u201cextreme\u201d version of our Inception module, with one spatial convolution per output channel of the 1x1 convolution. network design as early as 2014 [15] and has become more popular since its inclusion in the TensorFlow framework [1] in 2016. A depthwise separable convolution, commonly called", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 561 }, { "text": "\u201cseparable convolution\u201d in deep learning frameworks such as TensorFlow and Keras, consists in a depthwise convolution, i.e. a spatial convolution performed independently over each channel of an input, followed by a pointwise convolution, i.e. a 1x1 convolution, projecting the channels output by the", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 562 }, { "text": "depthwise convolution onto a new channel space. This is not to be confused with a spatially separable convolution, which is also commonly called \u201cseparable convolution\u201d in the image processing community. Two minor differences between and \u201cextreme\u201d version of an Inception module and a depthwise separable convolution", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 563 }, { "text": "would be: \u2022 The order of the operations: depthwise separable con- volutions as usually implemented (e.g. in TensorFlow) perform \ufb01rst channel-wise spatial convolution and then perform 1x1 convolution, whereas Inception performs the 1x1 convolution \ufb01rst. \u2022 The presence or absence of a non-linearity after the", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 564 }, { "text": "separable convolutions are usually implemented with- out non-linearities. We argue that the \ufb01rst difference is unimportant, in par- ticular because these operations are meant to be used in a stacked setting. The second difference might matter, and we investigate it in the experimental section (in particular see", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 565 }, { "text": "\ufb01gure 10). We also note that other intermediate formulations of In- ception modules that lie in between regular Inception mod- ules and depthwise separable convolutions are also possible: in effect, there is a discrete spectrum between regular convo- lutions and depthwise separable convolutions, parametrized", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 566 }, { "text": "by the number of independent channel-space segments used for performing spatial convolutions. A regular convolution (preceded by a 1x1 convolution), at one extreme of this spectrum, corresponds to the single-segment case; a depth- wise separable convolution corresponds to the other extreme where there is one segment per channel; Inception modules", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 567 }, { "text": "lie in between, dividing a few hundreds of channels into 3 or 4 segments. The properties of such intermediate modules appear not to have been explored yet. Having made these observations, we suggest that it may be possible to improve upon the Inception family of archi- tectures by replacing Inception modules with depthwise sep-", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 568 }, { "text": "arable convolutions, i.e. by building models that would be stacks of depthwise separable convolutions. This is made practical by the ef\ufb01cient depthwise convolution implementa- tion available in TensorFlow. In what follows, we present a convolutional neural network architecture based on this idea, with a similar number of parameters as Inception V3, and", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 569 }, { "text": "we evaluate its performance against Inception V3 on two large-scale image classi\ufb01cation task. 2. Prior work The present work relies heavily on prior efforts in the following areas: \u2022 Convolutional neural networks [10, 9, 25], in particular the VGG-16 architecture [18], which is schematically similar to our proposed architecture in a few respects.", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 570 }, { "text": "\u2022 The Inception architecture family of convolutional neu- ral networks [20, 7, 21, 19], which \ufb01rst demonstrated the advantages of factoring convolutions into multiple branches operating successively on channels and then on space. \u2022 Depthwise separable convolutions, which our proposed architecture is entirely based upon. While the use of spa-", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 571 }, { "text": "tially separable convolutions in neural networks has a long history, going back to at least 2012 [12] (but likely even earlier), the depthwise version is more recent. Lau- rent Sifre developed depthwise separable convolutions during an internship at Google Brain in 2013, and used them in AlexNet to obtain small gains in accuracy and", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 572 }, { "text": "large gains in convergence speed, as well as a signi\ufb01cant reduction in model size. An overview of his work was \ufb01rst made public in a presentation at ICLR 2014 [23]. Detailed experimental results are reported in Sifre\u2019s the- sis, section 6.2 [15]. This initial work on depthwise sep- arable convolutions was inspired by prior research from", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 573 }, { "text": "Sifre and Mallat on transformation-invariant scattering [16, 15]. Later, a depthwise separable convolution was used as the \ufb01rst layer of Inception V1 and Inception V2 [20, 7]. Within Google, Andrew Howard [6] has introduced ef\ufb01cient mobile models called MobileNets using depthwise separable convolutions. Jin et al. in", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 574 }, { "text": "2014 [8] and Wang et al. in 2016 [24] also did related work aiming at reducing the size and computational cost of convolutional neural networks using separable convolutions. Additionally, our work is only possible due to the inclusion of an ef\ufb01cient implementation of depthwise separable convolutions in the TensorFlow", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 575 }, { "text": "framework [1]. \u2022 Residual connections, introduced by He et al. in [4], which our proposed architecture uses extensively. 3. The Xception architecture We propose a convolutional neural network architecture based entirely on depthwise separable convolution layers. In effect, we make the following hypothesis: that the map-", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 576 }, { "text": "ping of cross-channels correlations and spatial correlations in the feature maps of convolutional neural networks can be entirely decoupled. Because this hypothesis is a stronger ver- sion of the hypothesis underlying the Inception architecture, we name our proposed architecture Xception, which stands", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 577 }, { "text": "for \u201cExtreme Inception\u201d. A complete description of the speci\ufb01cations of the net- work is given in \ufb01gure 5. The Xception architecture has 36 convolutional layers forming the feature extraction base of the network. In our experimental evaluation we will ex- clusively investigate image classi\ufb01cation and therefore our", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 578 }, { "text": "convolutional base will be followed by a logistic regression layer. Optionally one may insert fully-connected layers be- fore the logistic regression layer, which is explored in the experimental evaluation section (in particular, see \ufb01gures 7 and 8). The 36 convolutional layers are structured into 14 modules, all of which have linear residual connections", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 579 }, { "text": "around them, except for the \ufb01rst and last modules. In short, the Xception architecture is a linear stack of depthwise separable convolution layers with residual con- nections. This makes the architecture very easy to de\ufb01ne and modify; it takes only 30 to 40 lines of code using a high- level library such as Keras [2] or TensorFlow-Slim [17], not", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 580 }, { "text": "like architectures such as Inception V2 or V3 which are far more complex to de\ufb01ne. An open-source implementation of Xception using Keras and TensorFlow is provided as part of the Keras Applications module2, under the MIT license. 4. Experimental evaluation We choose to compare Xception to the Inception V3 ar-", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 581 }, { "text": "chitecture, due to their similarity of scale: Xception and Inception V3 have nearly the same number of parameters (table 3), and thus any performance gap could not be at- tributed to a difference in network capacity. We conduct our comparison on two image classi\ufb01cation tasks: one is the well-known 1000-class single-label classi\ufb01cation task on", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 582 }, { "text": "the ImageNet dataset [14], and the other is a 17,000-class multi-label classi\ufb01cation task on the large-scale JFT dataset. 4.1. The JFT dataset JFT is an internal Google dataset for large-scale image classi\ufb01cation dataset, \ufb01rst introduced by Hinton et al. in [5], which comprises over 350 million high-resolution images", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 583 }, { "text": "annotated with labels from a set of 17,000 classes. To eval- uate the performance of a model trained on JFT, we use an auxiliary dataset, FastEval14k. FastEval14k is a dataset of 14,000 images with dense annotations from about 6,000 classes (36.5 labels per im- age on average). On this dataset we evaluate performance", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 584 }, { "text": "using Mean Average Precision for top 100 predictions (MAP@100), and we weight the contribution of each class to MAP@100 with a score estimating how common (and therefore important) the class is among social media images. This evaluation procedure is meant to capture performance on frequently occurring labels from social media, which is", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 585 }, { "text": "crucial for production models at Google. 4.2. Optimization con\ufb01guration A different optimization con\ufb01guration was used for Ima- geNet and JFT: \u2022 On ImageNet: \u2013 Optimizer: SGD \u2013 Momentum: 0.9 \u2013 Initial learning rate: 0.045 \u2013 Learning rate decay: decay of rate 0.94 every 2 epochs \u2022 On JFT: \u2013 Optimizer: RMSprop [22]", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 586 }, { "text": "\u2013 Momentum: 0.9 \u2013 Initial learning rate: 0.001 2https://keras.io/applications/#xception \u2013 Learning rate decay: decay of rate 0.9 every 3,000,000 samples For both datasets, the same exact same optimization con- \ufb01guration was used for both Xception and Inception V3. Note that this con\ufb01guration was tuned for best performance", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 587 }, { "text": "with Inception V3; we did not attempt to tune optimization hyperparameters for Xception. Since the networks have dif- ferent training pro\ufb01les (\ufb01gure 6), this may be suboptimal, es- pecially on the ImageNet dataset, on which the optimization con\ufb01guration used had been carefully tuned for Inception V3.", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 588 }, { "text": "Additionally, all models were evaluated using Polyak averaging [13] at inference time. 4.3. Regularization con\ufb01guration \u2022 Weight decay: The Inception V3 model uses a weight decay (L2 regularization) rate of 4e \u22125, which has been carefully tuned for performance on ImageNet. We found this rate to be quite suboptimal for Xception", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 589 }, { "text": "and instead settled for 1e \u22125. We did not perform an extensive search for the optimal weight decay rate. The same weight decay rates were used both for the ImageNet experiments and the JFT experiments. \u2022 Dropout: For the ImageNet experiments, both models include a dropout layer of rate 0.5 before the logistic", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 590 }, { "text": "regression layer. For the JFT experiments, no dropout was included due to the large size of the dataset which made over\ufb01tting unlikely in any reasonable amount of time. \u2022 Auxiliary loss tower: The Inception V3 architecture may optionally include an auxiliary tower which back- propagates the classi\ufb01cation loss earlier in the network,", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 591 }, { "text": "serving as an additional regularization mechanism. For simplicity, we choose not to include this auxiliary tower in any of our models. 4.4. Training infrastructure All networks were implemented using the TensorFlow framework [1] and trained on 60 NVIDIA K80 GPUs each. For the ImageNet experiments, we used data parallelism", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 592 }, { "text": "with synchronous gradient descent to achieve the best classi- \ufb01cation performance, while for JFT we used asynchronous gradient descent so as to speed up training. The ImageNet experiments took approximately 3 days each, while the JFT experiments took over one month each. The JFT models were not trained to full convergence, which would have", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 593 }, { "text": "4.5. Comparison with Inception V3 4.5.1 Classi\ufb01cation performance All evaluations were run with a single crop of the inputs images and a single model. ImageNet results are reported on the validation set rather than the test set (i.e. on the non-blacklisted images from the validation set of ILSVRC 2012). JFT results are reported after 30 million iterations", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 594 }, { "text": "(one month of training) rather than after full convergence. Results are provided in table 1 and table 2, as well as \ufb01gure 6, \ufb01gure 7, \ufb01gure 8. On JFT, we tested both versions of our networks that did not include any fully-connected layers, and versions that included two fully-connected layers of 4096", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 595 }, { "text": "units each before the logistic regression layer. On ImageNet, Xception shows marginally better results than Inception V3. On JFT, Xception shows a 4.3% rel- ative improvement on the FastEval14k MAP@100 metric. We also note that Xception outperforms ImageNet results reported by He et al. for ResNet-50, ResNet-101 and ResNet-", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 596 }, { "text": "152 [4]. Table 1. Classi\ufb01cation performance comparison on ImageNet (sin- gle crop, single model). VGG-16 and ResNet-152 numbers are only included as a reminder. The version of Inception V3 being benchmarked does not include the auxiliary tower. Top-1 accuracy Top-5 accuracy VGG-16 0.715 0.901 ResNet-152", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 597 }, { "text": "0.770 0.933 Inception V3 0.782 0.941 Xception 0.790 0.945 The Xception architecture shows a much larger perfor- mance improvement on the JFT dataset compared to the ImageNet dataset. We believe this may be due to the fact that Inception V3 was developed with a focus on ImageNet and may thus be by design over-\ufb01t to this speci\ufb01c task. On", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 598 }, { "text": "ularization parameters) would yield signi\ufb01cant additional improvement. 4.5.2 Size and speed Table 3. Size and training speed comparison. Parameter count Steps/second Inception V3 23,626,728 31 Xception 22,855,952 28 In table 3 we compare the size and speed of Inception Figure 8. Training pro\ufb01le on JFT, with fully-connected layers", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 599 }, { "text": "V3 and Xception. Parameter count is reported on ImageNet (1000 classes, no fully-connected layers) and the number of training steps (gradient updates) per second is reported on ImageNet with 60 K80 GPUs running synchronous gradient descent. Both architectures have approximately the same size (within 3.5%), and Xception is marginally slower. We", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 600 }, { "text": "expect that engineering optimizations at the level of the depthwise convolution operations can make Xception faster than Inception V3 in the near future. The fact that both architectures have almost the same number of parameters indicates that the improvement seen on ImageNet and JFT does not come from added capacity but rather from a more", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 601 }, { "text": "ef\ufb01cient use of the model parameters. 4.6. Effect of the residual connections Figure 9. Training pro\ufb01le with and without residual connections. To quantify the bene\ufb01ts of residual connections in the Xception architecture, we benchmarked on ImageNet a mod- i\ufb01ed version of Xception that does not include any residual", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 602 }, { "text": "connections. Results are shown in \ufb01gure 9. Residual con- nections are clearly essential in helping with convergence, both in terms of speed and \ufb01nal classi\ufb01cation performance. However we will note that benchmarking the non-residual model with the same optimization con\ufb01guration as the resid- ual model may be uncharitable and that better optimization", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 603 }, { "text": "con\ufb01gurations might yield more competitive results. Additionally, let us note that this result merely shows the importance of residual connections for this speci\ufb01c architec- ture, and that residual connections are in no way required in order to build models that are stacks of depthwise sepa- rable convolutions. We also obtained excellent results with", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 604 }, { "text": "non-residual VGG-style models where all convolution layers were replaced with depthwise separable convolutions (with a depth multiplier of 1), superior to Inception V3 on JFT at equal parameter count. 4.7. Effect of an intermediate activation after point- wise convolutions Figure 10. Training pro\ufb01le with different activations between the", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 605 }, { "text": "depthwise and pointwise operations of the separable convolution layers. We mentioned earlier that the analogy between depth- wise separable convolutions and Inception modules suggests that depthwise separable convolutions should potentially in- clude a non-linearity between the depthwise and pointwise", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 606 }, { "text": "operations. In the experiments reported so far, no such non- linearity was included. However we also experimentally tested the inclusion of either ReLU or ELU [3] as intermedi- ate non-linearity. Results are reported on ImageNet in \ufb01gure 10, and show that the absence of any non-linearity leads to both faster convergence and better \ufb01nal performance.", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 607 }, { "text": "This is a remarkable observation, since Szegedy et al. re- port the opposite result in [21] for Inception modules. It may be that the depth of the intermediate feature spaces on which spatial convolutions are applied is critical to the usefulness of the non-linearity: for deep feature spaces (e.g. those", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 608 }, { "text": "found in Inception modules) the non-linearity is helpful, but for shallow ones (e.g. the 1-channel deep feature spaces of depthwise separable convolutions) it becomes harmful, possibly due to a loss of information. 5. Future directions We noted earlier the existence of a discrete spectrum be- tween regular convolutions and depthwise separable convo-", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 609 }, { "text": "lutions, parametrized by the number of independent channel- space segments used for performing spatial convolutions. In- ception modules are one point on this spectrum. We showed in our empirical evaluation that the extreme formulation of an Inception module, the depthwise separable convolution, may have advantages over regular a regular Inception mod-", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 610 }, { "text": "ule. However, there is no reason to believe that depthwise separable convolutions are optimal. It may be that intermedi- ate points on the spectrum, lying between regular Inception modules and depthwise separable convolutions, hold further advantages. This question is left for future investigation.", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 611 }, { "text": "6. Conclusions We showed how convolutions and depthwise separable convolutions lie at both extremes of a discrete spectrum, with Inception modules being an intermediate point in be- tween. This observation has led to us to propose replacing Inception modules with depthwise separable convolutions in", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 612 }, { "text": "neural computer vision architectures. We presented a novel architecture based on this idea, named Xception, which has a similar parameter count as Inception V3. Compared to Inception V3, Xception shows small gains in classi\ufb01cation performance on the ImageNet dataset and large gains on the JFT dataset. We expect depthwise separable convolutions", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 613 }, { "text": "to become a cornerstone of convolutional neural network architecture design in the future, since they offer similar properties as Inception modules, yet are as easy to use as regular convolution layers. References [1] M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghe-", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 614 }, { "text": "mawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Man\u00b4e, R. Monga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Van- houcke, V. Vasudevan, F. Vi\u00b4egas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng. Tensor-", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 615 }, { "text": "Flow: Large-scale machine learning on heterogeneous sys- tems, 2015. Software available from tensor\ufb02ow.org. [2] F. Chollet. Keras. https://github.com/fchollet/keras, 2015. [3] D.-A. Clevert, T. Unterthiner, and S. Hochreiter. Fast and accurate deep network learning by exponential linear units (elus). arXiv preprint arXiv:1511.07289, 2015.", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 616 }, { "text": "arXiv:1412.5474, 2014. [9] A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classi\ufb01cation with deep convolutional neural networks. In Advances in neural information processing systems, pages 1097\u20131105, 2012. [10] Y. LeCun, L. Jackel, L. Bottou, C. Cortes, J. S. Denker, H. Drucker, I. Guyon, U. Muller, E. Sackinger, P. Simard,", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 617 }, { "text": "et al. Learning algorithms for classi\ufb01cation: A comparison on handwritten digit recognition. Neural networks: the statistical mechanics perspective, 261:276, 1995. [11] M. Lin, Q. Chen, and S. Yan. Network in network. arXiv preprint arXiv:1312.4400, 2013. [12] F. Mamalet and C. Garcia. Simplifying ConvNets for Fast", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 618 }, { "text": "Learning. In International Conference on Arti\ufb01cial Neural Networks (ICANN 2012), pages 58\u201365. Springer, 2012. [13] B. T. Polyak and A. B. Juditsky. Acceleration of stochas- tic approximation by averaging. SIAM J. Control Optim., 30(4):838\u2013855, July 1992. [14] O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma,", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 619 }, { "text": "Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al. Ima- genet large scale visual recognition challenge. 2014. [15] L. Sifre. Rigid-motion scattering for image classi\ufb01cation, 2014. Ph.D. thesis. [16] L. Sifre and S. Mallat. Rotation, scaling and deformation invariant scattering for texture discrimination. In 2013 IEEE", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 620 }, { "text": "Conference on Computer Vision and Pattern Recognition, Portland, OR, USA, June 23-28, 2013, pages 1233\u20131240, 2013. [17] N. Silberman and S. Guadarrama. Tf-slim, 2016. [18] K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 621 }, { "text": "Computer Vision and Pattern Recognition, pages 1\u20139, 2015. [21] C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna. Rethinking the inception architecture for computer vision. arXiv preprint arXiv:1512.00567, 2015. [22] T. Tieleman and G. Hinton. Divide the gradient by a run- ning average of its recent magnitude. COURSERA: Neural", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 622 }, { "text": "Networks for Machine Learning, 4, 2012. Accessed: 2015- 11-05. [23] V. Vanhoucke. Learning visual representations at scale. ICLR, 2014. [24] M. Wang, B. Liu, and H. Foroosh. Factorized convolutional neural networks. arXiv preprint arXiv:1608.04337, 2016. [25] M. D. Zeiler and R. Fergus. Visualizing and understanding", "source": "Xception", "year": 2017, "url": "https://arxiv.org/abs/1610.02357", "id": 623 }, { "text": "arXiv:2004.10448v1 [cs.CV] 22 Apr 2020 DeepFake Detection by Analyzing Convolutional Traces Luca Guarnera University of Catania - iCTLab Catania, Italy luca.guarnera@unict.it Oliver Giudice University of Catania Catania, Italy giudice@dmi.unict.it Sebastiano Battiato University of Catania - iCTLab", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 624 }, { "text": "Catania, Italy battiato@dmi.unict.it Abstract The Deepfake phenomenon has become very popular nowadays thanks to the possibility to create incredibly real- istic images using deep learning tools, based mainly on ad- hoc Generative Adversarial Networks (GAN). In this work we focus on the analysis of Deepfakes of human faces with", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 625 }, { "text": "the objective of creating a new detection method able to de- tect a forensics trace hidden in images: a sort of \ufb01nger- print left in the image generation process. The proposed technique, by means of an Expectation Maximization (EM) algorithm, extracts a set of local features speci\ufb01cally ad- dressed to model the underlying convolutional generative", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 626 }, { "text": "process. Ad-hoc validation has been employed through experimental tests with naive classi\ufb01ers on \ufb01ve different architectures (GDWCT, STARGAN, ATTGAN, STYLEGAN, STYLEGAN2) against the CELEBA dataset as ground-truth for non-fakes. Results demonstrated the effectiveness of the technique in distinguishing the different architectures and", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 627 }, { "text": "the corresponding generation process. 1. Introduction One of the phenomena that is rapidly growing is the well- known Deepfake: the possibility to automatically generate and/or alter/swap a person\u2019s face in images and videos us- ing algorithms based on Deep Learning technology. It is possible to generate excellent results by creating new mul-", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 628 }, { "text": "timedia contents that cannot be easily recognized as real or fake by human eye. Then, the term Deepfake refers to all those multimedia contents synthetically altered or created by means of machine learning generative models. Various examples of Deepfake, involving celebrities, are easily discoverable on the web: the insertion of Nicholas", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 629 }, { "text": "Cage 1 in movies where he did not act like \u201cFight Club\u201d and \u201cThe Matrix\u201d or the impressive video in which Jim Carrey 2 plays Shining in place of Jack Nicholson. Other 1https://www.youtube.com/watch?v=-yQxsIWO2ic 2https://www.youtube.com/watch?v=Dx59bskG8dc more worrying examples are the video of ex US Presi-", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 630 }, { "text": "dent Barack Obama (Figure 1(a)), created by Buzzfeed 3 in collaboration with Monkeypaw Studios, or the video in which Mark Zuckerberg 4 (Figure 1(b)) claims a series of statements about his platform ability to steal users\u2019 data. Even in Italy, in September 2019, the satirical TV program \u201cStriscia La Notizia\u201d 5 showed a video of the ex-premier", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 631 }, { "text": "Matteo Renzi talking about his colleagues in a \u201cnot so re- spectful\u201d way (Figure 1 (c)). Indeed, Deepfakes may have serious repercussions on the authenticity of the news spread by the mass-media while representing a new threat for pol- itics, companies and individual privacy. In this dangerous scenario, tools are needed to unmask the Deepfakes or just", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 632 }, { "text": "detect them. Several big companies have decided to take action against this phenomenon: Google has created a database of fake videos [36] to support researchers who are devel- oping new techniques to detect them, while Facebook and Microsoft have launched the Deepfake Detection Challenge initiative 6.", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 633 }, { "text": "In this paper a new Deepfake detection method will be introduced focused on images representing human faces. At \ufb01rst an Expectation Maximization (EM) algorithm [29], extracts a set of local features speci\ufb01cally addressed to model the convolutional traces that could be found in im- ages. Then, naive classi\ufb01ers were trained to discriminate", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 634 }, { "text": "between authentic images and images generated by the \ufb01ve most realistic architectures as today (GDWCT, STARGAN, ATTGAN, STYLEGAN, STYLEGAN2). Experimental re- sults demonstrated that the information modelled by EM is related to the speci\ufb01c architecture that generated the im- age thus giving the overall detection solution explainability,", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 635 }, { "text": "being also of great value for forensic investigations (e.g., camera model identi\ufb01cation techniques of image forensics). Moreover, a multitude of experiments will be presented not only to demonstrated the effectiveness of the technique but 3https://www.youtube.com/watch?v=cQ54GDm1eL0 4https://www.youtube.com/watch?v=NbedWhzx1rs", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 636 }, { "text": "also to demonstrate to be un-comparable with state-of-the- art: tests were carried out on an almost-in-the-wild dataset with images generated by \ufb01ve different techniques with dif- ferent image sizes. As today, all proposed technique work with speci\ufb01c image sizes and against at most one GAN tech- nique.", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 637 }, { "text": "The remainder of this paper is organized as follows: Sec- tion 2 presents some Deepfake generation and detection methods. The proposed detection technique is explained in Section 3 as regards the feature extraction phase while the classi\ufb01cation phase and experimental results are reported in Section 4. Finally, Section 5 concludes the paper with in-", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 638 }, { "text": "sights for future works. 2. Related Works Deepfakes are generally created by techniques based on Generative Adversarial Networks (GANs) \ufb01rstly introduced by Goodfellow et al. [14]. Authors proposed a new frame- work for estimating generative models via an adversarial mode in which two models simultaneously train: a gen-", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 639 }, { "text": "erative model G, that captures the data distribution, and a discriminative model D, able to estimate the probabil- ity that a sample comes from the training data rather than from G. The training procedure for G is to maximize the probability of D making a mistake thus resulting to a min- max two-player game. Mathematically, the generator ac-", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 640 }, { "text": "cepts a random input z with density pz and returns an output x = G(z, \u0398g) according to a certain probability distribution pg (\u0398g represent the parameters of the generative model). The discriminator, D(x, \u0398d) computes the probability that x comes from the distribution of training data pdata (\u0398d represents the parameters of the discriminative model). The", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 641 }, { "text": "overall objective is to obtain a generator, after the training phase, which is a good estimator of pdata. When this hap- pens, the discriminator is \u201cdeceived\u201d and will no longer be able to distinguish the samples from pdata and pg; there- fore pg will follow the targeted probability distribution, i.e.", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 642 }, { "text": "pdata. Figure 2 shows a simpli\ufb01ed description of a GAN framework. In the case of Deepfakes, G can be thought as a team of counterfeiters trying to produce fake currency, while D stands to the police, trying to detect the malicious activity. G and D can be implemented as any kind of gen- erative model, in particular when deep neural networks are", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 643 }, { "text": "employed results become extremely accurate. Through re- cent years, many GAN architectures were proposed for dif- ferent applications e.g., image to image translation [45], im- age super resolution [24], image completion [17], and text- to-image generation [35]. 2.1. Deepfake Generation Techniques An overview on Media forensics with particular focus on", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 644 }, { "text": "Deepfakes has been recently proposed in [41]. STARGAN is a method capable of performing image- to-image translations on multiple domains using a single model. Proposed by Choi et al. [6] was trained on two dif- ferent types of face datasets: CELEBA [27] containing 40 labels related to facial attributes such as hair color, gender", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 645 }, { "text": "and age, and RaFD dataset [22] containing 8 labels corre- sponding to different types of facial expressions (\u201chappy\u201d, \u201csad\u201d, etc.). Given a random label as input, such as hair color, facial expression, etc., STARGAN is able to per- form an image-to-image translation operation. Results have been compared with other existing methods [26, 30, 45] and", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 646 }, { "text": "showed how STARGAN manages to generate images of su- perior visual quality. Style Generative Adversarial Network, namely STYLE- GAN [19], changed the generator model of STARGAN by means of mapping points in latent space to an intermediate latent space which controls the style output at each point of", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 647 }, { "text": "strates to achieve better results. Thus, STYLEGAN is capa- ble not only of generating impressively photorealistic and high-quality photos of faces, but also offers control param- eters in terms of the overall style the generated image at different levels of detail. While being able to create real- istic pseudo-portraits, small details might reveal the fake-", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 648 }, { "text": "ness of generated images. To correct those imperfections in STYLEGAN, Karras et al. made some improvements to the generator (including re-designed normalization, multi- resolution, and regularization methods) proposing STYLE- GAN2 [20]. Instead of imposing constraints on latent representation, He et al. [15], proposed a new technique called ATTGAN", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 649 }, { "text": "in which an attribute classi\ufb01cation constraint is applied to the generated image, in order to guarantee only the correct modi\ufb01cations of the desired attributes. The authors used CELEBA [27] and LFW [16] datasets, and performed var- ious tests comparing ATTGAN with VAE/GAN [23], Ic- GAN [30] and STARGAN [6], Fader Networks [21], Shen", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 650 }, { "text": "et al. [37] and CycleGAN [45]. Achieved results showed that ATTGAN exceeds the state of the art on the realistic modi\ufb01cation of facial attributes. The latter style transfer approach worth to be mentioned is the work of Cho et al. [5], where they propose a group- wise deep whitening-and coloring method (GDWCT) for", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 651 }, { "text": "a better styling capacity. They used CELEBA [27], Art- works [45], cat2dog [25], Ink pen and watercolor classes from Behance Artistic Media (BAM) [43], and Yosemite datasets [45] as dataset. GDWCT has been compared with various cutting-edge methods in image translation and style transfer improving not only computational ef\ufb01ciency but", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 652 }, { "text": "also quality of generated images. In this paper, the \ufb01ve most famous and effective archi- tectures in state-of-the-art for face Deepfakes were taken into account: STARGAN [6], STYLEGAN [19], STYLE- GAN2 [20], ATTGAN [15] and GDWCT [5]. As described above, they are different in goals and structure. Table 1 re-", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 653 }, { "text": "sumes the differences of the techniques in terms of image size, dataset and type of input, goal and architecture struc- ture. 2.2. Deepfake detection methods Being able to understand if an image is the result of a generative Neural Network process turns out to be a compli- cated problem, even for human eyes. However, the problem", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 654 }, { "text": "of authenticating an image (or speci\ufb01cally a digital image) is not new [2, 31, 38]. Many works try to reconstruct the history of an image[13]; others try to identify the anoma- lies, such as the study on the analysis of interpolation effects through CFA (Color Filtering Array) [32], analyzing com- pression parameters [3, 11, 12], etc. Given the peculiarity", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 655 }, { "text": "of Deepfakes, state-of-the-art image analysis methods tend to fail and more re\ufb01ned ones are needed. Thanks to a new discriminator that uses \u201ccontrastive loss\u201d it is possible to \ufb01nd the typical characteristics of the synthesized images generated by different GANs and therefore detect such fake images by means of a classi-", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 656 }, { "text": "\ufb01er. Rossler et al. [36] proposed an automated bench- mark for fake detection, based mainly on four manip- ulation methods: two computer graphics-based methods (Face2Face [40], FaceSwap 7) and 2 learning-based ap- proaches (DeepFakes 8, NeuralTextures [39]). They ad- dressed the problem of fake detection as a binary classi\ufb01-", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 657 }, { "text": "cation problem for each frame of manipulated videos, con- sidering different techniques present in the state of the art [1, 4, 7, 8, 10, 34]. Zhang et al. [44] proposed a method to classify Deep- fakes considering the spectra of the frequency domain as input. The authors proposed a GAN simulation framework,", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 658 }, { "text": "called AutoGAN, in order to emulate the process commonly shared by popular GAN models. Results obtained by the au- thors achieved very good performances in terms of binary classi\ufb01cation between authentic and fake images. Also Du- rall et al. [9] presented a method for Deepfakes detection based on the analysis in the frequency domain. The authors", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 659 }, { "text": "combined high-resolution authentic face images from dif- ferent public datasets (CELEBA-HQ data set [18], Flickr- Faces-HQ data set [19]) with fakes (100K Faces project 9, this person does not exist 10), creating a new dataset called Faces-HQ. By means of naive classi\ufb01ers they obtained good results in terms of overall accuracy.", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 660 }, { "text": "Differently from described approaches, in this paper the possibility to capture the underlying traces of a possible Deepfake is investigated by employing a sort of reverse en- gineering of the last computational layer of a given GAN architecture. This method will give explainability to the predictions of Deepfakes being of great value for forensic", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 661 }, { "text": "investigations: not only it is able to classify an image as fake but also can predict the most probable technique used for generation being in this way similar to camera model de- tection in image forensics analysis [2]. The underlying idea of the technique is to \ufb01nd the main periodic components (e.g. transpose computational layer) on generated images.", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 662 }, { "text": "A similar strategy was proposed some time ago in a sem- inal paper of Popescu et al. [32] devoted to point out the presence of digital forgeries in CFA interpolated images. Another difference from state-of-the-art is the working sce- nario: the proposed technique demonstrates to achieve good results in a almost-in-the-wild scenario with images gener-", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 663 }, { "text": "Method Number of images generated Size Data input to the network Goal of the network Kernel size of the latest Convolution Layer GDWCT [5] 3369 216x216 CELEBA Improves the styling capability 4x4 STARGAN [6] 5648 256x256 CELEBA Image-to-image translations on multiple domains using a single model 7x7", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 664 }, { "text": "ATTGAN [15] 6005 256x256 CELEBA Transfer of face attributes with classi\ufb01cation constraints 4x4 STYLEGAN [19] 9999 1024x1024 CELEBA-HQ FFHQ Transfer semantic content from a source domain to a target domain characterized by a different style 3x3 STYLEGAN2 [20] 3000 1024x1024 FFHQ Transfer semantic content from a", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 665 }, { "text": "source domain to a target domain characterized by a different style 3x3 Table 1. Details of Deepfake GAN architectures employed for analysis. For each one is reported: all images generated, the generated image sizes, the original input used to train the neural network, the goal of the network and the kernel size of last convolutional layer.", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 666 }, { "text": "3. Extracting Convolutional Traces The most common and effective technical solutions able to generate Deepfakes are the Generative Adversarial Net- works speci\ufb01cally deep ones. For all the techniques de- scribed before, the generator G is composed of Transpose Convolution layers [33]. In Neural Networks like CNNs,", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 667 }, { "text": "Convolution operations apply a \ufb01lter, namely kernel, to the input multidimensional array. After each convolution layer a pooling operation is needed to reduce output dimensional size w.r.t. input. On the other hand, in generative models the Transpose Convolution Layers are employed. They also ap- ply kernels to input but they act inversely in order to obtain", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 668 }, { "text": "an output larger but proportional to the input dimensions. The starting idea of the proposed approach is that local correlation of pixels in Deepfakes are dependent exclusively on the operations performed by all the layers present in the GAN which generate it; speci\ufb01cally the (latter) transpose convolution layers. In order to \ufb01nd these trace, unsuper-", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 669 }, { "text": "vised machine learning techniques were taken into account. Indeed, different unsupervised learning techniques aim at creating clusters containing instances of the input dataset with high similarity between instances of the same cluster while having high dissimilarity between instances belong- ing to different clusters. These clusters can represent the", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 670 }, { "text": "\u201chidden\u201d structure of the dataset analyzed. Therefore, the clustering technique must estimate which are the parame- ters of the distributions that most likely generated the train- ing samples. Based on this principle, an Expectation Max- imization (EM) algorithm [29] was employed in order to de\ufb01ne a conceptual mathematical model able to capture the", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 671 }, { "text": "pixel correlation of the images (e.g. spatially). The result of EM is a feature vector representing the structure of the Transpose Convolution Layers employed during the gener- ation of the image, encoding in some sense is such images if a Deepfake or not. The initial goal is to extract a description, from input", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 672 }, { "text": "image I, able to numerically represent the local correlations between each pixel in a neighbourhood. This can be done by means of convolution with a kernel k of N \u00d7 N size: I[x, y] = \u03b1 X s,t=\u2212\u03b1 ks,t \u2217I[x + s, y + t] (1) In Equation 1, the value of the pixel I[x, y] is computed considering a neighborhood of size N \u00d7N of the input data.", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 673 }, { "text": "It is clear that the new estimated information I[x, y] mainly depends on the kernel used in the convolution operation, which establishes a mathematical relationship between the pixels. For this reason, our goal is to de\ufb01ne a vector k of size N \u00d7 N able to capture this hidden and implicit relationship", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 674 }, { "text": "which characterize of forensic trace we want to exploit. Let\u2019s assume that the element I[x, y] belongs to one of the following models: \u2022 M1: when the element I[x, y] satis\ufb01es Equation 1; \u2022 M2: otherwise. The EM algorithm is employed with its two different steps: 1. Expectation step: computes the (density of) probabil-", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 675 }, { "text": "and unknown variance and M2 uniform. In the Expectation step, the Bayes rule that I[x, y] belongs to the model M1 is computed as follows: Pr{I[x, y] \u2208M1 | I[x, y]} = = Pr{I[x, y] | I[x, y] \u2208M1} \u2217Pr{I[x, y] \u2208M1} 2P i=1 Pr{I[x, y] | I[x, y] \u2208Mi} \u2217Pr{I[x, y] \u2208Mi} (2) where the probability distribution of M1 which repre-", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 676 }, { "text": "sents the probability of observing a sample I[x, y], knowing that it was generated by the model M1 is: Pr{I[x, y] | I[x, y] \u2208M1} = 1 \u03c3 \u221a 2\u03c0 e\u2212(R[x,y])2 2\u03c32 (3) where R[x, y] = I[x, y] \u2212 \u03b1 X s,t=\u2212\u03b1 ks,tI[x + s, y + t] (4) . The variance value \u03c32, which is still unknown, is then estimated in the Maximization step. Once de\ufb01ned if I[x, y]", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 677 }, { "text": "belongs to model M1 (or M2), the values of the vector \u20d7k are estimated using least squares method, minimizing the following: E(\u20d7k) = X x,y w[x, y] I[x, y]\u2212 \u03b1 X s,t=\u2212\u03b1 ks,tI[x + s, y + t] !2 (5) where w \u2261Pr{I[x, y] \u2208M1 | I[x, y]} (2). This error function (5) can be minimized by computing the gradient of", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 678 }, { "text": "vector \u20d7k. The update of ki,j is carried out by computing the partial derivative of (5) as follows: \u2202E \u2202ki,j = 0 (6) Hence, the following linear equations system is obtained: \u03b1 X s,t=\u2212\u03b1 ks,t X x,y w[x, y]I[x + i, y + j]I[x + s, y + t] ! = = X x,y w[x, y]I[x + i, y + j]I[x, y] (7) The two steps of the EM algorithm are iteratively re-", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 679 }, { "text": "peated. A pseudo-code description is provided in Algorithm Algorithm 1: Expectation-Maximization Algorithm Data: Image I Result: \u20d7k Initialize N //Kernel size Initialize \u03c30 Set \u20d7k random of size NxN Set R, P, W matrices with 0 values of the same size as I Set p0 as 1/size of the range of values of I", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 680 }, { "text": "for n = 1; n < 100 n+ = 1 do //Expectation Step for \u2200values in I do R[x, y] = I[x, y]\u2212 \u03b1P s,t=\u2212\u03b1 ks,tI[x + s, y + t] P[x, y] = 1 \u03c3n \u221a 2\u03c0e \u2212R[x,y] 2\u03c32n W[x, y] = P [x,y] P [x,y]+p0 //Maximization Step Calculate k(n+1) s,t as shown in the formula 7 1:Expectation-Maximization. The algorithm is applied to", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 681 }, { "text": "each channel of the input image (RGB color space). The obtained feature vector \u20d7k, has dimensions dependent to parameter \u03b1. Note that the element k0,0 will always be set equal to 0 (k0,0 = 0). Thus, for example, if a kernel k with 3 \u00d7 3 size is employed, the resulting \u20d7k will be a vector of 24 elements (since the values k0,0 are excluded). This is", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 682 }, { "text": "obtained by concatenating the features extracted from each of the three RGB channels. The computational complexity of the EM algorithm can be estimated to be linear in d (the number of characteristics of the input data taken into consideration), n (the number of objects) and t (the number of iterations).", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 683 }, { "text": "4. Classi\ufb01cation Phase and Results Six datasets of images were taken into account for train- ing and testing purposes: one containing only authentic face images of celebrities (CELEBA), and the others contain- ing DeepFakes generated by \ufb01ve different GANs (STAR- GAN, STYLEGAN, STYLEGAN2, GDWCT, ATTGAN).", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 684 }, { "text": "For STYLEGAN and STYLEGAN2, images were down- loaded from STYLEGAN 11 and STYLEGAN2 12 respec- tively; while STARGAN, ATTGAN and GDWCT were em- ployed in inference mode to generate their respective image datasets. An overview of the DeepFake data generated for each GAN is reported in the Table 1. The EM algorithm, as described in previous Section, was", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 685 }, { "text": "increasing sizes (3, 4, 5 and 7) 13. The obtained feature vec- tor was employed as input of different naive classi\ufb01ers (K- NN, SVM and LDA) with different tasks: (i) discriminat- ing authentic image from one speci\ufb01c GAN and (ii) dis- criminating authentic images from Deepfakes. The overall classi\ufb01cation pipeline of the proposed approach is brie\ufb02y", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 686 }, { "text": "summarized in Figure 3. Let\u2019s \ufb01rst analyse the discrim- inative power of the extracted feature vector in order to distinguish authentic images (CELEBA) from each of the considered GANs (CELEBA Vs STARGAN, CELEBA Vs STYLEGAN, CELEBA Vs STYLEGAN2, CELEBA Vs ATTGAN, CELEBA Vs GDWCT). Figure 4 shows a vis-", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 687 }, { "text": "ible representation by means of t-SNE [28]: in which it is possible to notice, how some categories of networks that create Deepfake can be \u201clinearly\u201d separable from authen- tic samples. However in most case the separation is utterly clear. Classi\ufb01cation tests were carried out on the obtained fea- ture vectors with, as expected from what seen from t-SNE", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 688 }, { "text": "representation, excellent results. All the classi\ufb01cation re- sults are reported in Table 2. In particular, it is possible to note that: \u2022 CELEBA Vs ATTGAN the maximum classi\ufb01cation accuracy of 92.67%, was obtained with KNN - K = 3, and kernel size of 3x3. \u2022 CELEBA Vs GDWCT: the maximum classi\ufb01cation", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 689 }, { "text": "13Typical kernel size used by the latest Transpose Convolution Layers (which have a fundamental role in the creation of the Deepfake images) of the different GAN architectures accuracy of 88.40%, was obtained with KNN - K = 3,5,7, and kernel size of 3x3. \u2022 CELEBA Vs STARGAN: the maximum classi\ufb01ca- tion accuracy of 93.17%, was obtained with linear", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 690 }, { "text": "SVM, and kernel size of 7x7. \u2022 CELEBA Vs STYLEGAN: the maximum classi\ufb01ca- tion accuracy of 99.65%, was obtained with KNN - K = 3,5,7,9, and kernel size of 4x4. \u2022 CELEBA Vs STYLEGAN2: the maximum classi\ufb01- cation accuracy of 99.81%, was obtained with linear SVM, and kernel size of 4x4. The kernel size used by convolution layers in the neu-", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 691 }, { "text": "ral networks represents one of the elements to identify the forensic trace that we are looking for. Table 1 shows the kernel size (and other information) of the neural networks that we have taken into account for our experiments. As described above, the structure of the GAN plays a fundamental role in the Deepfakes detection, in particular", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 692 }, { "text": "for what regards the generator structure. Considering the images from STYLEGAN and STYLEGAN2, it is possible to distinguish them, as the authors of the STYLEGAN2 ar- chitecture have only updated parts of the generator in order to remove some imperfections of STYLEGAN. This fur- ther con\ufb01rms the hypothesis, since even a slight modi\ufb01ca-", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 693 }, { "text": "CELEBA Vs ATTGAN CELEBA Vs GDWCT CELEBA Vs STARGAN CELEBA Vs STYLEGAN CELEBA Vs STYLEGAN2 Kernel Size Kernel Size Kernel Size Kernel Size Kernel Size 3x3 4x4 5x5 7x7 3x3 4x4 5x5 7x7 3x3 4x4 5x5 7x7 3x3 4x4 5x5 7x7 3x3 4x4 5x5 7x7 3-NN 92.67 86.50 84.50 85.33 88.40 73.17 73.00 74.33 90.50 89.00 88.67", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 694 }, { "text": "85.17 93.00 99.65 98.26 99.55 96.99 99.61 98.75 97.77 5-NN 92.00 86.50 84.83 86.17 88.40 75.67 74.17 76.67 88.83 88.83 88.17 85.00 93.00 99.65 98.26 99.32 97.39 99.61 98.21 97.55 7-NN 91.00 87.67 85.33 85.67 88.40 76.67 71.33 78.67 89.33 89.17 88.00 84.83 93.50 99.65 98.07 99.09 97.39 99.42 98.21 97.55", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 695 }, { "text": "9-NN 90.83 87.67 84.83 86.50 87.70 76.83 71.17 79.00 89.33 89.17 87.50 84.67 92.83 99.65 98.07 99.32 97.19 99.42 98.39 97.10 11-NN 91.00 86.83 85.33 85.83 88.05 76.67 72.83 77.00 89.17 88.67 86.67 83.50 93.17 99.48 98.07 99.32 96.99 99.42 97.85 97.10 13-NN 91.00 87.17 84.50 85.33 87.87 75.33 73.50 77.17", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 696 }, { "text": "88.33 89.33 87.50 83.50 93.50 99.48 98.07 99.09 97.39 99.22 97.67 97.10 SVM 90.50 89.67 90.33 87.00 87.35 76.50 79.00 80.50 90.00 88.50 88.83 93.17 92.00 98.96 99.42 98.41 96.99 99.81 99.46 97.77 LDA 89.50 88.50 89.50 87.17 87.52 76.00 79.33 81.67 89.67 87.83 88.83 90.00 92.50 99.31 98.84 99.09 96.79", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 697 }, { "text": "99.61 99.10 97.77 Table 2. Overall accuracy between CELEBA vs. each one of the considered GANs. Results are presented w.r.t. all the different kernel sizes (3x3, 4x4, 5x5, 7x7) and with different classi\ufb01ers: KNN, with k \u2208{3, 5, 7, 9, 11, 13}; Linear SVM, Linear Discriminant Analysis (LDA). CELEBA Vs DeepNetworks", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 698 }, { "text": "Kernel Size 3x3 4x4 5x5 7x7 3-NN 89.96 84.90 80.76 82.69 5-NN 90.22 86.63 82.48 82.77 7-NN 89.57 87.12 82.48 84.27 9-NN 89.51 86.73 83.31 84.27 11-NN 89.25 87.21 83.69 83.97 13-NN 89.57 87.31 84.20 83.45 SVMLinear 88.02 88.75 86.05 85.85 SVMsigmoid 86.08 72.60 83.38 63.66 SVMrbf 89.77 89.71 86.24 87.43", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 699 }, { "text": "SVMPoly 82.51 86.06 84.65 86.61 LDA 87.56 88.65 86.11 85.48 Table 3. Overall accuracy between CELEBA with all Deep Neu- ral Network, with different kernel size (3x3, 4x4, 5x5, 7x7 - ob- tained through the EM algorithm) and with different classi\ufb01ers used: KNN, with k \u2208{3, 5, 7, 9, 11, 13}; SVM (linear, sigmoid,", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 700 }, { "text": "rbf, polynomial), Linear Discriminant Analysis (LDA). STYLEGAN Vs STYLEGAN2 Kernel Size 3x3 4x4 5x5 7x7 3-NN 89.36 83.57 90.51 87.24 5-NN 89.56 86.41 89.87 85.52 7-NN 89.16 85.40 90.93 87.59 9-NN 88.55 83.98 89.87 87.93 11-NN 88.35 83.37 90.30 87.24 13-NN 89.36 82.76 89.66 87.93 SVM 91.77 95.13 99.16", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 701 }, { "text": "99.31 LDA 91.16 94.52 98.73 98.28 Table 4. Overall accuracy between STYLEGAN and STYLE- GAN2, with all the different kernel size (3x3, 4x4, 5x5, 7x7 - obtained through the EM algorithm) and with different employed classi\ufb01ers: KNN, with k \u2208{3, 5, 7, 9, 11, 13}; Linear SVM, Lin- ear Discriminant Analysis (LDA).", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 702 }, { "text": "Finally, another type of classi\ufb01cation was the compar- ison between CELEBA original images and all the im- ages generated with all the networks as a binary classi\ufb01- cation problem. In this test, a further analysis of the two- dimensional t-SNE was carried out. Figure 5) shows that, in this case, samples cannot be linearly separated.", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 703 }, { "text": "For this reason, other non-linear classi\ufb01ers were taken into ac- count reaching a maximum accuracy of 90.22% (with KNN, K=5), with kernel employed in the EM of size 3\u00d73. Table 3 shows the obtained results in the binary classi\ufb01cation task. Many additional experiments were carried out to fur- therly demonstrate the effectiveness of the extracted fea-", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 704 }, { "text": "ture vector as a descriptor of the hidden convolutional trace. Speci\ufb01cally results w.r.t. classi\ufb01cation tests between dif- ferent combinations of GANs are described furtherly con- ferming the robustness of the technique. Also other t-SNE representations are provided and can be found at the follow- ing address https://iplab.dmi.unict.it/mfs/DeepFake/.", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 705 }, { "text": "Finally, it is worth to point out that during the research activity a deep neural network technique was employed to detect Deepfakes on the datasets described above. Tests car- ried out with VGG-1614 on both spatial and frequency do- main of images achieved a best result of 53% of accuracy in the binary classi\ufb01cation task showing that a deep learning", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 706 }, { "text": "approach is not able to extract what the proposed approach was able to. Our results are similar in terms of overall per- formance by experiments exploited in Wang et al. [42] that is actually able to reach very high results by simply using a discriminator trained on one family of GANs and using it to infer if images are real or generated from other types of", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 707 }, { "text": "GANs. 5. Conclusions and future works The \ufb01nal result of our study to counter the Deepfake phe- nomenon was the creation of a new detection method based on features extracted through the EM algorithm. The under- lying \ufb01ngerprint has been proven to be effective to discrim- inate between images generated by recent GANs architec-", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 708 }, { "text": "tures speci\ufb01cally devoted to generate realistic people\u2019s face. Some more works will be devoted to investigate the role of the kernel dimensions. Also the possibility to extend such methodology to video\u2019s analysis and/or evaluate the robust- ness with respect to standard image editing (e.g. photomet-", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 709 }, { "text": "kernel size): CELEBA Vs DeepNetworks. will be considered. In general one of the key aspect will the possibility to adapt the method in situations on the \u201cwild\u201d without any a-priori knowledge of the generation process. Acknowledgement This research was supported by iCTLab s.r.l. - Spin- off of University of Catania (https://www.ictlab.srl), which", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 710 }, { "text": "provided domain expertise and computational power that greatly assisted the activity. References [1] D. Afchar, V. Nozick, J. Yamagishi, and I. Echizen. Mesonet: a compact facial video forgery detection network. In 2018 IEEE International Workshop on In- formation Forensics and Security (WIFS), pages 1\u20137.", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 711 }, { "text": "IEEE, 2018. 3 [2] S. Battiato, O. Giudice, and A. Paratore. Multimedia forensics: discovering the history of multimedia con- tents. In Proceedings of the 17th International Con- ference on Computer Systems and Technologies 2016, pages 5\u201316. ACM, 2016. 3 [3] S. Battiato and G. Messina. Digital forgery estimation", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 712 }, { "text": "into DCT domain: a critical analysis. In Proceedings of the First ACM Workshop on Multimedia in Foren- sics, pages 37\u201342, 2009. 3 [4] B. Bayar and M. C Stamm. A deep learning approach to universal image manipulation detection using a new convolutional layer. In Proceedings of the 4th ACM Workshop on Information Hiding and Multimedia Se-", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 713 }, { "text": "curity, pages 5\u201310, 2016. 3 [5] W. Cho, S. Choi, D. K. Park, I. Shin, and J. Choo. Image-to-image translation via group-wise deep whitening-and-coloring transformation. In Pro- ceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 10639\u201310647, 2019. 3, 4 [6] Y. Choi, M. Choi, M. Kim, J. Ha, S. Kim, and J. Choo.", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 714 }, { "text": "Stargan: Uni\ufb01ed generative adversarial networks for multi-domain image-to-image translation. In Proceed- ings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 8789\u20138797, 2018. 2, 3, 4 [7] F. Chollet. Xception: Deep learning with depthwise separable convolutions. In Proceedings of the IEEE", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 715 }, { "text": "Unmasking deepfakes with simple features. arXiv preprint arXiv:1911.00686, 2019. 3 [10] J. Fridrich and J. Kodovsky. Rich models for steganal- ysis of digital images. IEEE Transactions on Informa- tion Forensics and Security, 7(3):868\u2013882, 2012. 3 [11] F. Galvan, G. Puglisi, A. R. Bruna, and S. Battiato.", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 716 }, { "text": "First quantization matrix estimation from double com- pressed JPEG images. IEEE Transactions on Informa- tion Forensics and Security, 9(8):1299\u20131310, 2014. 3 [12] O. Giudice, F. Guarnera, A. Paratore, and S. Battiato. 1-D DCT domain analysis for JPEG double compres- sion detection. In Proceeedings of International Con-", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 717 }, { "text": "ference on Image Analysis and Processing, pages 716\u2013 726. Springer, 2019. 3 [13] O. Giudice, A. Paratore, M. Moltisanti, and S. Bat- tiato. A Classi\ufb01cation Engine for Image Ballistics of Social Data, pages 625\u2013636. Springer International Publishing, 2017. 3 [14] I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D.", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 718 }, { "text": "Warde-Farley, S. Ozair, A. Courville, and Y. Bengio. Generative adversarial nets. In Advances in Neural Information Processing Systems, pages 2672\u20132680, 2014. 2 [15] Z. He, W. Zuo, M. Kan, S. Shan, and X. Chen. Attgan: Facial attribute editing by only changing what you want. IEEE Transactions on Image Processing,", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 719 }, { "text": "28(11):5464\u20135478, 2019. 3, 4 [16] G. B Huang, M. Mattar, T. Berg, and E. Learned- Miller. Labeled faces in the wild: A database forstudy- ing face recognition in unconstrained environments. In Workshop on Faces in \u2019Real-Life\u2019 Images: De- tection, Alignment, and Recognition, Oct 2008, Mar- seille, France. inria-00321923, 2008. 3", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 720 }, { "text": "3 [19] T. Karras, S. Laine, and T. Aila. A style-based gener- ator architecture for generative adversarial networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4401\u20134410, 2019. 2, 3, 4 [20] T. Karras, S. Laine, M. Aittala, J. Hellsten, J. Lehti- nen, and T. Aila. Analyzing and improving the image", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 721 }, { "text": "quality of stylegan. arXiv preprint arXiv:1912.04958, 2019. 3, 4 [21] G. Lample, N. Zeghidour, N. Usunier, A. Bordes, L. Denoyer, and M. Ranzato. Fader networks: Manip- ulating images by sliding attributes. In Advances in Neural Information Processing Systems, pages 5967\u2013 5976, 2017. 3 [22] O. Langner, R. Dotsch, G. Bijlstra, D. HJ Wigboldus,", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 722 }, { "text": "S. T Hawk, and AD Van Knippenberg. Presentation and validation of the radboud faces database. Cogni- tion and emotion, 24(8):1377\u20131388, 2010. 2 [23] A. Boesen Lindbo Larsen, S. Kaae S\u00f8nderby, H. Larochelle, and O. Winther. Autoencoding beyond pixels using a learned similarity metric. arXiv preprint", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 723 }, { "text": "arXiv:1512.09300, 2015. 3 [24] C. Ledig, L. Theis, F. Husz\u00b4ar, J. Caballero, A. Cun- ningham, A. Acosta, A. Aitken, A. Tejani, J. Totz, Z. Wang, et al. Photo-realistic single image super- resolution using a generative adversarial network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4681\u20134690,", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 724 }, { "text": "2017. 2 [25] H. Lee, H. Tseng, J. Huang, M. Singh, and M. Yang. Diverse image-to-image translation via disentangled representations. In Proceedings of the European Con- ference on Computer Vision (ECCV), pages 35\u201351, 2018. 3 [26] M. Li, W. Zuo, and D. Zhang. Deep identity- aware transfer of facial attributes.", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 725 }, { "text": "arXiv preprint arXiv:1610.05586, 2016. 2 [27] Z. Liu, P. Luo, X. Wang, and X. Tang. Deep learning face attributes in the wild. In Proceedings of the IEEE International Conference on Computer Vision, pages 3730\u20133738, 2015. 2, 3 [28] L. van der Maaten and G. Hinton. Visualizing data using t-sne. Journal of Machine Learning Research,", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 726 }, { "text": "9(Nov):2579\u20132605, 2008. 6 [29] T. K Moon. The expectation-maximization algo- rithm. IEEE Signal Processing Magazine, 13(6):47\u2013 60, 1996. 1, 4 [30] G. Perarnau, J. Van De Weijer, B. Raducanu, and J. M \u00b4Alvarez. Invertible conditional GANs for image edit- ing. arXiv preprint arXiv:1611.06355, 2016. 2, 3", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 727 }, { "text": "arXiv:1511.06434, 2015. 4 [34] N. Rahmouni, V. Nozick, J. Yamagishi, and I. Echizen. Distinguishing computer graphics from natural im- ages using convolution neural networks. In 2017 IEEE Workshop on Information Forensics and Secu- rity (WIFS), pages 1\u20136. IEEE, 2017. 3 [35] S. Reed, Z. Akata, X. Yan, L. Logeswaran, B. Schiele,", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 728 }, { "text": "and H. Lee. Generative adversarial text to image syn- thesis. arXiv preprint arXiv:1605.05396, 2016. 2 [36] A. Rossler, D. Cozzolino, L. Verdoliva, C. Riess, J. Thies, and M. Nie\u00dfner. Faceforensics++: Learning to detect manipulated facial images. In Proceedings of the IEEE International Conference on Computer Vi-", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 729 }, { "text": "sion, pages 1\u201311, 2019. 1, 3 [37] W. Shen and R. Liu. Learning residual images for face attribute manipulation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recogni- tion, pages 4030\u20134038, 2017. 3 [38] M. C. Stamm, M. Wu, and K. J. R. Liu. Information forensics: An overview of the \ufb01rst decade. IEEE Ac-", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 730 }, { "text": "cess, 1:167\u2013200, 2013. 3 [39] J. Thies, M. Zollh\u00a8ofer, and M. Nie\u00dfner. Deferred neu- ral rendering: Image synthesis using neural textures. ACM Transactions on Graphics (TOG), 38(4):1\u201312, 2019. 3 [40] J. Thies, M. Zollhofer, M. Stamminger, C. Theobalt, and M. Nie\u00dfner. Face2face: Real-time face capture", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 731 }, { "text": "and reenactment of RGB videos. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2387\u20132395, 2016. 3 [41] L. Verdoliva. Media forensics and deepfakes: an overview. arXiv preprint arXiv:2001.06564, 2020. 2 [42] S. Wang, O. Wang, R. Zhang, A. Owens, and A. Efros. Cnn-generated images are surprisingly easy to", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 732 }, { "text": "spot...for now. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition,, 2020. 7 [43] M. J Wilber, C. Fang, H. Jin, A. Hertzmann, J. Col- lomosse, and S. Belongie. Bam! the behance artistic media dataset for recognition beyond photography. In Proceedings of the IEEE International Conference on", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 733 }, { "text": "Computer Vision, pages 1202\u20131211, 2017. 3 [44] X. Zhang, S. Karaman, and S. Chang. Detecting and simulating artifacts in GAN fake images. arXiv preprint arXiv:1907.06515, 2019. 3 [45] J. Zhu, T. Park, P. Isola, and A. Efros. Unpaired image-to-image translation using cycle-consistent ad- versarial networks. In Proceedings of the IEEE In-", "source": "Watch Your Up-Convolution", "year": 2020, "url": "https://arxiv.org/abs/2004.10448", "id": 734 }, { "text": "Deepfake Generation and Detection: A Benchmark and Survey GAN PEI\u2217, East China Normal University, China JIANGNING ZHANG\u2217, Zhejiang University, China MENGHAN HU\u2020, East China Normal University, China ZHENYU ZHANG, Nanjing University, China CHENGJIE WANG, Youtu Lab, Tencent, China YUNSHENG WU, Youtu Lab, Tencent, China", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 735 }, { "text": "GUANGTAO ZHAI, Shanghai Jiao Tong University, China JIAN YANG, Nanjing University, China DACHENG TAO, Nanyang Technological University, Singapore Deepfake technology aims to synthesize highly realistic facial images and videos, with broad application potential in entertainment, film production, and digital human modeling. Deep learning has driven major", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 736 }, { "text": "progress in generative modeling, from VAEs and GANs to the recent rise of diffusion models. The latter have sparked a renewed wave of research through their superior generation quality. In addition to deepfake generation, corresponding detection technologies continuously evolve to regulate the potential misuse of", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 737 }, { "text": "deepfakes, such as privacy invasion and phishing attacks. This survey comprehensively reviews the latest developments in deepfake generation and detection, summarizing and analyzing current state-of-the-arts in this rapidly evolving field. First, we unify task definitions, comprehensively introduce datasets and metrics,", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 738 }, { "text": "and summarize the underlying technologies. Then, we review the development of several related sub-fields and examine four representative deepfake research fields: face swapping, face reenactment, talking-face generation, and facial attribute editing, as well as forgery detection. Subsequently, we benchmark representative methods", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 739 }, { "text": "on widely adopted datasets to provide a comprehensive and up-to-date evaluation of the most influential published works. Finally, we discuss the key challenges and outline future research directions for the field. We closely follow the latest developments in this project. Additional Key Words and Phrases: Deepfake Generation, Face Swapping, Face Reenactment, Talking Face", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 740 }, { "text": "Generation, Facial Attribute Editing, Forgery Detection, Survey. 1 Introduction Artificial Intelligence Generated Content (AIGC) garners considerable attention [1] in academia and industry. Deepfake generation, as one of the important technologies in the generative domain, gains significant attention due to its ability to create highly realistic facial media content. This technique", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 741 }, { "text": "transitions from traditional graphics-based methods to deep learning-based approaches. Early meth- ods employ advanced Variational Autoencoder [2\u20134] (VAE) and Generative Adversarial Network (GAN) [5, 6] techniques, enabling seemingly realistic image generation, but their performance is still unsatisfactory, which limits practical applications. Recently, the diffusion structure [7\u20139] has", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 742 }, { "text": "\u2217Both authors contributed equally to this research \u2020Corresponding author Authors\u2019 addresses: Gan Pei, 52295904023@stu.ecnu.edu.cn, East China Normal University, China; Jiangning Zhang, 186368@zju.edu.cn, Zhejiang University, China; Menghan Hu, mhhu@ce.ecnu.edu.cn, East China Normal University, China;", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 743 }, { "text": "Zhenyu Zhang, zhenyuzhang@nju.edu.cn, Nanjing University, China; Chengjie Wang, jasoncjwang@tencent.com, Youtu Lab, Tencent, China; Yunsheng Wu, simonwu@tencent.com, Youtu Lab, Tencent, China; Guangtao Zhai, zhaiguangtao@ sjtu.edu.cn, Shanghai Jiao Tong University, China; Jian Yang, csjianyang@gmail.com, Nanjing University, China; Dacheng", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 744 }, { "text": "2 Pei and Zhang, et al. greatly enhanced the generation capability of images and videos. Benefiting from this new wave of research, deepfake technology demonstrates potential value for practical applications and can generate content indistinguishable from real ones, which has further attracted attention and is", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 745 }, { "text": "widely applied in numerous fields [10], including entertainment, movie production, etc. Deepfake generation can generally be divided into four mainstream research fields: 1) Face swapping [11\u201313] is dedicated to executing identity exchanges between two person images; 2) Face reenactment [14, 15] emphasizes transferring source movements and poses; 3) Talking face", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 746 }, { "text": "generation [16, 17] focuses on achieving natural matching of mouth movements to textual content in character generation, and 4) Facial attribute editing [18\u201320] aims to modify specific facial attributes of the target image. The development of related foundational technologies has gradually shifted from single forward GAN models [5, 21] to multi-step diffusion models [7, 22, 23] with higher", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 747 }, { "text": "quality generation capabilities, and the generated content has also gradually transitioned from single-frame images to temporal video modeling [8]. In addition, NeRF [24, 25] has been frequently incorporated into modeling to improve multi-view consistency capabilities [26, 27]. While enjoying the novelty and convenience of this technology, its unethical use has raised", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 748 }, { "text": "serious societal and ethical concerns. In practice, deepfake content has been used to generate non-consensual explicit videos targeting individuals. Moreover, there have been instances where videos of deceased public figures were synthetically recreated and disseminated without the consent of their families, causing emotional distress and raising serious moral and legal questions.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 749 }, { "text": "These incidents illustrate substantial risks of privacy invasion, identity impersonation, and large- scale dissemination of misleading or harmful content. Consequently, there is an urgent need for effective forgery detection systems to mitigate malicious misuse and safeguard information security [28, 29]. From the earliest handcrafted feature-based methods [30, 31] to deep learning-", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 750 }, { "text": "based methods [32, 33], and the recent hybrid detection techniques [34], forgery detection has undergone substantial technological advancements along with the development of generative technologies. The data modality has also transitioned from the spatial and frequent domains [35, 36] to the more challenging temporal domain [37, 38]. Considering that current generative technologies", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 751 }, { "text": "have a higher level of interest, develop faster, and can generate indistinguishable content from reality [39], corresponding detection technologies need continuous evolution. Overall, despite notable progress in both directions, existing methods still face limitations in visual authenticity and generative accuracy under challenging scenarios [40]. These issues continue", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 752 }, { "text": "to draw research attention and raise questions about their practical deployment. However, prior surveys cover only parts of the deepfake landscape and overlook emerging technologies [1, 39, 41], particularly diffusion-based image and video generation. This survey provides a comprehensive overview of these areas and related sub-fields while also tracking the latest developments.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 753 }, { "text": "\u2022 Contribution. In this survey, we comprehensively explore the key technologies and latest advancements in Deepfakes generation and forgery detection. We first unify the task definitions (Sec. 2.1), provide a comprehensive comparison of datasets and metrics (Sec. 2.3), and discuss the development of related technologies. Then, we investigate four mainstream deepfake fields, as", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 754 }, { "text": "well as forgery detection (Sec. 3). We also analyze the benchmarks and settings for each domain, thoroughly evaluating the latest and influential works published in top-tier conferences/journals (Sec. 4), especially recent diffusion-based approaches. Additionally, we discuss closely related fields, including head swapping, face super-resolution, face reconstruction, face inpainting, body", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 755 }, { "text": "animation, portrait style transfer, makeup transfer, and adversarial sample detection. \u2022 Scope. This survey primarily focuses on four mainstream face-related tasks and forgery detection. We also cover some related domain tasks in Sec. 2.4 and detail specific popular sub-tasks in Sec. 3.3. Considering the large number of articles (including published and preprints), we mainly include", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 756 }, { "text": "Deepfake Generation and Detection: A Benchmark and Survey 3 \u00a71 Introduction \u00a72 Background \u00a73 Deepfake Inspections \u00a72.1 Problem Definition \u00a72.2 Technological History and Roadmap \u00a72.3 Datasets, Metrics, and Losses \u00a72.4 Related Research Domains \u00a72.5 Applications of Deepfake Detection \u00a72.6 Ethical and Societal Considerations", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 757 }, { "text": "\u00a73.1.1 Face Swapping Traditional Graphics GAN-based Diffusion-based Alternative Techniques \u00a73.3 Specific Related Domains Face Super-resolution Portrait Style Transfer Body Animation Makeup Transfer \u00a73.1.2 Face Reenactment 3DMM-based Landmark Matching Feature Decoupling Self-supervised Learning \u00a73.1.3 Talking Face Generation", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 758 }, { "text": "Audio/Text Driven Multimodal Conditioned Diffusion-based 3D-model Technologies \u00a73.1.4 Facial Attribute Editing Comprehensive Editing Irrelevant-attribute Retained Diffusion-based Text Driven \u00a73.2 Forgery Detection Space Domain Time Domain Frequency Domain Data Driven \u00a74 Benchmark \u00a74.1 Metrics \u00a74.2 Benchmark Protocol", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 759 }, { "text": "\u00a74.3 Main Results on Deepfake Generation \u00a74.4 Main Results on Forgery Detection \u00a75 Prospects \u00a76 Conclusion Face Swapping Face Reenactment Talking Face Generation Facial Attribute Editing Forgery Detection Discussion Fig. 1. Time diagram that reflects the survey pipeline. Zoom in for a better holistic perception of this work.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 760 }, { "text": "surveys. Sha et al. [1] only discuss character generation while we cover a more comprehensive range of tasks. Compared to works [39, 41, 42], our study encompasses a broader range of technical models, particularly the more powerful diffusion-based methods. Additionally, we thoroughly discuss the related sub-fields of deepfake generation and detection.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 761 }, { "text": "\u2022 Survey Methods. To conduct an effective literature review, we systematically searched major academic databases, including IEEE Xplore, ACM Digital Library, Springer, and ScienceDirect, and screened technical articles that met our inclusion criteria. Specifically, we focused on studies published between 2020 and 2025, while grouping research published prior to 2020 into a separate", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 762 }, { "text": "category. For each research topic, we combined relevant technical keywords and application- oriented keywords to perform targeted searches. The technical keywords included Generative Adversarial Networks (GANs), GAN variants, diffusion models, autoencoders, and Neural Radiance Fields (NeRF), among others. The application keywords included deepfake generation, deepfake", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 763 }, { "text": "detection, face swapping, face replacement, face reenactment, talking face, talking head, and forgery detection. The retrieved results were manually deduplicated and categorized according to application domains. Representative works for each research topic were selected based on publication venue quality, citation count, and technical novelty.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 764 }, { "text": "\u2022 Survey Pipeline. Fig. 1 illustrates the overall structure of this survey. Sec. 2 introduces the essential background, including task definitions, datasets, evaluation metrics, and related research areas. It also outlines key applications of deepfake technologies and discusses their ethical impli-", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 765 }, { "text": "cations. Sec. 3 presents a technical review of the four major deepfake tasks, organized from the perspective of methodological categorization and technological evolution. Sec. 4 summarizes and compares the performance of representative methods on widely used benchmarks to provide a fair and comprehensive evaluation. Sec. 5 critically examines remaining challenges and highlights", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 766 }, { "text": "potential directions for future research. Finally, Sec. 6 provides a concise summary of the survey. 2 Background In this section, we first introduce the conceptual definitions of the discussed mainstream fields. Fig. 2 illustrates the intuitive objectives for each task and shows the distinctions among tasks in terms", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 767 }, { "text": "of manipulated facial components. Then, we review the developmental history of commonly used neural networks, highlighting several representative ones. Next, we summarize popular datasets, metrics, and loss functions. Subsequently, we comprehensively discuss several relevant domains. Finally, we introduce some application and discussion about there ethical considerations.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 768 }, { "text": "2.1 Problem Definition \u2022 Unified Formulation of Studied Problems. Fig. 2 intuitively displays the various deepfakes generation and detection tasks studied in this paper. For the former, different tasks can essentially be expressed as controlled content generation problems under specific conditions, such as images,", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 769 }, { "text": "4 Pei and Zhang, et al. We have now reached the tipping point of generative AI \ud835\udc3c! \ud835\udc36 Face Swapping Face Reenactment Talking Face Generation Age Editing Gender Change Mouth Editing \u2026 Facial Attribute Editing Forgery Detection \u2205# Real or Fake \u2205$ \ud835\udc3c% \u00a73.1.1 \u00a73.1.2 \u00a73.1.3 \u00a73.1.4 \u00a73.2 Identity Preservation", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 770 }, { "text": "Head Pose / Movement Temporal Continuity Multiple Attribute Editing \ud835\udc46% Fig. 2. Top: Illustration of different deepfake generation and detection tasks that are discussed in this survey. Bottom: Specific facial attribute modification of each task. Data from NVIDIA Keynote at COMPUTEX 2023. audio, text, specific attributes, etc.. Given the target image \ud835\udc3c\ud835\udc61to be manipulated and the condition", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 771 }, { "text": "information \ud835\udc36= {Image, Audio, Text, . . . }, the content generation process can be represented as: \ud835\udc3c\ud835\udc5c= \ud835\udf53\ud835\udc6e(\ud835\udc3c\ud835\udc61,\ud835\udc36), (1) where \ud835\udf19\ud835\udc3aabstracts the specific generation network and \ud835\udc3c\ud835\udc5c= {\ud835\udc3c0 \ud835\udc61, \ud835\udc3c1 \ud835\udc61, . . . , \ud835\udc3c\ud835\udc41\u22121 \ud835\udc61 } represents generated contents. \ud835\udc41is the total frame number for the generated video, which is set to 1 by default). The latter", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 772 }, { "text": "task can be viewed as an image-level or pixel-level classification problem as practical application needs, which can be represented as: \ud835\udc46\ud835\udc5c= \ud835\udf53\ud835\udc6b(\ud835\udc3c\ud835\udc5c), (2) where \ud835\udf19\ud835\udc37abstracts the detection network and \ud835\udc46\ud835\udc5crepresents the fake score for generated content \ud835\udc3c\ud835\udc5c. \u2022 Face Swapping. This task replaces the identity of the target face \ud835\udc3c\ud835\udc61with that of the source face", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 773 }, { "text": "\ud835\udc3c\ud835\udc60, while maintaining target-specific, ID-irrelevant attributes such as skin tone and expressions. \u2022 Face Reenactment. This task transfers the facial movements from a driving image or video to a target image \ud835\udc3c\ud835\udc61, while keeping the target\u2019s identity and attributes unchanged. It commonly relies on facial motion capture techniques, including tracking or deep-learning-based motion prediction.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 774 }, { "text": "\u2022 Talking Face Generation. This task generates a talking video \ud835\udc3c\ud835\udc5cfor the character in a target image \ud835\udc3c\ud835\udc61, driven by text, audio, video, or multimodal inputs. The output should accurately reflect the driving information, including lip motion, facial pose, emotions, and spoken content. \u2022 Facial Attribute Editing. This task modifies semantic attributes of a target face \ud835\udc3c\ud835\udc61(e.g., age,", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 775 }, { "text": "expression, or skin tone) in a controlled way according to user intent. Methods fall into single- attribute and multi-attribute editing, with the latter is the main focus of this survey. \u2022 Forgery Detection. This task detects and localizes tampering or forged regions in images or videos using an anomaly score \ud835\udc46\ud835\udc5c. It is importance to information security and multimedia forensics.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 776 }, { "text": "2.2 Technological History and Roadmap \u2022 Generative Framework. VAEs [2\u20134], GANs [5, 6, 21], and Diffusion [22, 23, 43] have played pivotal roles in the developmental history of generative models. 1) VAE [2] emerges in 2013, altering the relationship between latent features and linear mappings in autoencoder latent spaces. It introduces feature distributions like the Gaussian distribution and", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 777 }, { "text": "Deepfake Generation and Detection: A Benchmark and Survey 5 2019-2023 After 2023 Before 2019 Pix2Pix (ICCV\u201917) GAN (arxiv\u201914) CVAE (NIPS\u201915) VAE (Stat\u201914) Diffusion (ICML\u201915) CGAN (ICCV\u201917) CVAE-GAN (ICCV\u201917) VQ-VAE (NIPS\u201917) RealForensics (CVPR\u201922) M2TR (ICMR\u201922) RECCE (CVPR\u201922) SBIs (CVPR\u201922) MRL", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 778 }, { "text": "(T INF FOREN SEC\u201923 NoiseDF (AAAI\u201923) Guo et al. (T INF FOREN SEC\u201923) Guo et al. (ICCV\u201923) AVoiD-DF (T INF FOREN SEC\u201923 Feng et al. (CVPR\u201923) Lgrad (CVPR\u201923) HiFi-Net (CVPR\u201923) DDIM (ICML\u201920) StyleGAN3 (NIPS\u201921) StyleGAN (CVPR\u201919) VQ-VAE2 (NIPS\u201919) LDM (CVPR\u201922) DDPM (NIPS\u201920) StyleGAN2 (CVPR\u201922) CycleGAN", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 779 }, { "text": "(CVPR\u201919) Animatediff (arXiv\u201923) SVD (arXiv\u201923) Collaborative-Diff (CVPR\u201923) LDGM (CVPR\u201923) AVoiD-DF (T INF FOREN SEC\u201923 Feng et al. (CVPR\u201923) Lgrad (CVPR\u201923) HiFi-Net (CVPR\u201923) Cat4d (CVPR\u201925) SinDiffusion (TPAMI\u201925) EMDM (ECCV\u201924) Lumiere (SIGGRAPH\u201924) Fig. 3. Development timeline of three mainstream", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 780 }, { "text": "generative modals, i.e., VAE, GAN, and Diffusion. 0 20 40 60 80 100 120 140 160 Before 2020 2020 2021 2022 2023 2024 2025 Face Swap Face Reenactment Talking Face Generation Face Attribute Editing Forgery Detection Fig. 4. Works summarization on different directions per year. Data is obtained on 2025/11/20.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 781 }, { "text": "under specific conditions, CVAE [3] introduces conditional input. VQ-VAE [4] introduces the concept of vector quantization to improve the learning of latent representations. 2) GANs [21] generate realistic data through an adversarial process between two networks: a generator and a discriminator. This relationship can be likened to a boxing match in which the", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 782 }, { "text": "generator acts as the challenger attempting to produce increasingly convincing fake samples, while the discriminator serves as the defending champion distinguishing real from generated content. As the two opponents improve through continual competition, the quality of generated outputs steadily increases. Built upon this adversarial paradigm, GAN-based methods have rapidly", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 783 }, { "text": "expanded and remain central to many deepfake tasks. Extensions can be viewed as introducing new strategies to this \u201cmatch.\u201d CGAN [44] conditions both networks on auxiliary information, making the competition more guided. Pix2Pix [45] specializes the framework for image-to-image translation, akin to training the challenger for specific techniques. StyleGAN [5] introduces style-based controls", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 784 }, { "text": "for fine-grained feature manipulation, and StyleGAN2 [6] further stabilizes training to enhance quality and controllability. Hybrid models such as CVAE-GAN [46] enrich the generator\u2019s latent modeling by integrating VAE-based representations with adversarial learning. 3) Diffusion models [43] treat data generation as a gradual denoising process. Starting from near-", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 785 }, { "text": "random noise, the model learns to reverse the corruption step by step, eventually recovering a clean and realistic image. DDPM [22] popularizes this paradigm by achieving outstanding generative performance, particularly on large-scale and high-resolution images. LDM [23] improves efficiency and flexibility by performing the denoising process in a compact latent space, similar to restoring", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 786 }, { "text": "the compressed negative of an image rather than the full-resolution photograph. SVD [7] fine-tunes a base text-to-video model for image-to-video conversion, akin to extending a single frame into a coherent sequence. AnimateDiff [8] attaches a motion modeling module to a frozen text-to-image backbone and trains it on short video clips.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 787 }, { "text": "\u2022 Discriminative Neural Network. Convolutional Neural Networks (CNNs) [47\u201352] have played a pivotal role in the history of deep learning. LeNet [47], as the pioneer of CNNs, showcased the charm of machine learning. AlexNet [48] and ResNet [49] made deep learning feasible. Recently, ConvNeXt [50] has achieved excellent results surpassing those of Swin-Transformer [53]. The", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 788 }, { "text": "Transformer architecture initially proposed [54] in 2017. The core idea involves using self-attention mechanisms to capture dependencies between different positions in the input sequence, enabling global modeling of sequences. ViT [55] demonstrates that using Transformer in the field of computer vision can still achieve excellent performance. In addition, Swin-Transformer [53] addresses the", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 789 }, { "text": "limitations of Transformer in handling high-resolution image processing tasks. Swin-Transformer V2 [56] further improves the model\u2019s efficiency and the resolution of manageable inputs. \u2022 Neural Radiance Field. NeRF, introduced in 2020 [24], leverages volume rendering and implicit neural fields to represent and reconstruct the geometry and illumination of 3D scenes [25]. Com-", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 790 }, { "text": "6 Pei and Zhang, et al. Table 1. Overview of commonly used datasets. Orange-marked are selected to evaluate different methods. Dataset Type Scale Highlight Deepfake Generation LFW [63] Image 10K Facial images captured under various lighting conditions, poses, expressions, and occlusions at 250\u00d7250 resolution.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 791 }, { "text": "CelebA [64] Image 200K The dataset includes over 200K facial images from more than 10K individuals, each with 40 attribute labels. VGGFace [66] Image 2600K A super-large-scale facial dataset involving a staggering 26K participants, encompassing a total of 2600K facial images. VoxCeleb1 [69] Video 100K voices", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 792 }, { "text": "A large scale audio-visual dataset of human speech, the audio includes noisy background interference. CelebA-HQ [65] Image 30K A high-resolution facial dataset consisting of 30K face images, each with a resolution of 1024\u00d71024 resolution. VGGFace2 [67] Image 3000K A large-scale facial dataset, expanded with greater diversity in terms of ethnicity and pose compared to VGGFace.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 793 }, { "text": "VoxCeleb2 [70] Video 2K hours A dataset that is five times larger in scale than VoxCeleb1 and improves racial diversity. FFHQ [5] Image 70K The dataset with over 70K high-resolution (1024\u00d71024) facial images, showcasing diverse ethnicity, age, and backgrounds. MEAD [71] Video 40 hours An emotional audiovisual dataset provides facial expression information during conversations with various emotional labels.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 794 }, { "text": "CelebV-HQ [74] Video 35K videos The dataset\u2019s video clips have a resolution of more than 512\u00d7512 resolution and are annotated with rich attribute labels. Forgery Detection DeepfakeTIMIT [79] Audio-Video 640 videos The dataset is evenly divided into two versions: LQ (64\u00d764) and HQ (128\u00d7128), with all videos using face swapping forgery.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 795 }, { "text": "FF++ [80] Video 6K videos Comprising 1K original videos and manipulated videos generated using five different forgery methods. DFDCP [82] Audio-Video 5K videos The preliminary dataset for The Deepfake Detection Challenge includes two face-swapping methods. DFD [86] Video 3K videos The dataset comprises 3K deepfake videos generated using five forgery methods.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 796 }, { "text": "Deeperforensics [81] Video 60K videos The videos feature faces with diverse skin tones, and rich environmental diversity was considered during the filming process. DFDC [83] Audio-Video 128K videos The official dataset for The Deepfake Detection Challenge and contain a substantial amount of interference information.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 797 }, { "text": "Celeb-DF [84] Video 1K videos Comprising 408 genuine videos from diverse age groups, ethnicities, and genders, along with 795 DeepFake videos. Celeb-DFv2 [84] Video 6K videos An expanded version of Celeb-DFv1, this dataset not only increases in quantity but also the diversity. FakeAVCeleb [85] Audio-Video", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 798 }, { "text": "20K videos A novel audio-visual multimodal deepfake detection dataset, deepfake videos generated using four forgery methods. In addition, Some notable works [61, 62] combining NeRF as a supplement to 3D information and generation models are particularly prominent at present. \u2022 Work Summary. The evolution of mainstream generative models is depicted chronologically in", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 799 }, { "text": "Fig. 3. This survey delves into four categories of generation tasks along with the forgery detection task, and the publication years distribution of the surveyed articles is shown in Fig. 4. 2.3 Datasets, Metrics, and Losses \u2022 Dataset. Given the various datasets in surveyed fields, we use numerical labels to save post-textual", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 800 }, { "text": "space. 1) Commonly used deepfake generation datasets include LFW [63], CelebA [64], CelebA- HQ [65], VGGFace [66], VGGFace2 [67], FFHQ [5], Multi-PIE [68], VoxCeleb1 [69], VoxCeleb2 [70], MEAD [71], MM CelebA-HQ [72], CelebAText-HQ [73], CelebV-HQ [74], TalkingHead-1KH [75], LRS2 [76], LRS3 [77], etc. 2) Commonly used forgery detection datasets include UADFV [78], Deep-", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 801 }, { "text": "fakeTIMIT [79], FF++ [80], Deeperforensics-1.0 [81], DFDCP [82], DFDC [83], Celeb-DF [84], Celeb- DFv2 [84], FakeAVCeleb [85], DFD [86], WildDeepfake [87], KoDF [88], UADFV [89], Deephy [90], DF-Platter [91], etc. We summarize popular datasets in Tab. 1 . \u2022 Metric. 1) For deepfake generation tasks, commonly used metrics include: Peak Signal-to-Noise", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 802 }, { "text": "Ratio (PSNR) [92], Structured Similarity (SSIM) [93], Learned Perceptual Image Patch Similarity (LPIPS) [94], Fr\u00e9chet Inception Distance (FID) [95], Kernel Inception Distance (KID) [96], Cosine Similarity (CSIM) [75], Identity Retrieval Rate (ID Ret) [97], Expression Error [98], Pose Error [99], Landmark Distance (LMD) around the mouths [100], Lip-sync Confidence (LSE-C) [101], Lip-sync", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 803 }, { "text": "Distance (LSE-D) [101], etc. 2) For forgery generation commonly uses: Area Under the ROC Curve (AUC) [102], Accuracy (ACC) [103], Equal Error Rate (EER) [104], Average Precision (AP) [105], F1-Score [106], etc. Detailed definitions are explained in Sec. 4.1. \u2022 Loss Function. VAE-based approaches generally employ reconstruction loss and KL divergence", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 804 }, { "text": "loss [2]. Commonly used reconstruction loss functions include Mean Squared Error, Cross-Entropy, LPIPS [94], and perceptual [107] losses. GAN-based methods further introduce adversarial loss [21] to increase image authenticity, while diffusion-based works introduce denoising loss function [22]. 2.4", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 805 }, { "text": "Related Research Domains \u2022 Head Swapping. This task replaces the entire head region of a target image, including facial contours and hairstyle, with that of a source image [108]. Despite the larger replacement area, only identity-related attributes are transferred, while other facial attributes remain unchanged. Recently,", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 806 }, { "text": "diffusion-based methods [109] have emerged and demonstrated promising performance. \u2022 Face Super-resolution. This task enhances low-resolution images to produce high-resolution outputs [110, 111]. It is closely related to many deepfake sub-tasks. Early methods in Face Swap- ping [112, 113] and Talking Face Generation [114, 115] often suffered from low-resolution outputs", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 807 }, { "text": "Deepfake Generation and Detection: A Benchmark and Survey 7 in synthesized images and videos. This limitation has been alleviated by integrating face super- resolution modules into generative pipelines [16, 116], significantly improving visual quality. Tech- nically, face super-resolution approaches can be categorized into CNNs [117\u2013119], GANs [120, 121],", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 808 }, { "text": "reinforcement learning [122], and ensemble learning [123]. \u2022 Face Reconstruction. This task refers to reconstructing the three-dimensional facial structure of an individual from one or multiple 2D images [124, 125]. Facial reconstruction often serves as an intermediate step in various deepfake sub-tasks. In Face Swapping and Face Reenactment, 3DMM are", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 809 }, { "text": "widely used to recover facial parameters, enabling controllable identity or expression manipulation. Reconstructing a 3D facial model also helps mitigate artifacts that occur in synthesized videos under large pose variations. Technical approaches to face reconstruction include 3DMM [126, 127], epipolar-geometry [128], one-shot learning [129, 130], shadow shape reconstruction [131, 132],", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 810 }, { "text": "and hybrid learning-based reconstruction [133, 134]. \u2022 Face Inpainting. This task aims to reconstruct missing regions in face images caused by external factors such as occlusion and lighting while preserving facial texture information is crucial in this process [135]. This task is a crucial sub-task of image inpainting, and the current methods are mostly", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 811 }, { "text": "based on deep learning that can be roughly divided into two categories: GAN based [136, 137] and Diffusion based [138, 139]. \u2022 Body Animation. This task aims to alter the entire bodily pose while unchanging the overall body information [140]. The goal is to achieve a modification of the target image\u2019s entire body", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 812 }, { "text": "posture using an optional driving image or video, aligning the body of the target image with the information from the driving signal. The mainstream implementation path for body animation is based on GANs [141, 142], and Diffusion [143\u2013146]. \u2022 Portrait Style Transfer. This task aims to reconstruct the style of a target image to match that", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 813 }, { "text": "of a source image by learning the stylistic features of the source image [147, 148]. The goal is to preserve the content information of the target image while adapting its style to that of the source image [149]. Common applications include image cross-domain style transfer, such as transforming real face images into animated face styles [150, 151]. Methods based on GANs [152\u2013154] and", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 814 }, { "text": "Diffusion [155, 156] have achieved high-quality performance in this task. \u2022 Makeup Transfer. This task aims to achieve style transfer learning from a source image to a target image [157, 158]. Existing models have achieved initial success in applying and removing makeup on target images [158\u2013161], allowing for quantitative control over the intensity of makeup.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 815 }, { "text": "However, they perform poorly in transferring extreme styles [162\u2013164]. Existing mainstream methods are based on GANs [157, 165]. \u2022 Adversarial Sample Detection. This task focuses on identifying whether the input data is an adversarial sample [166]. If recognized as such, the model can refuse to provide services for it, such", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 816 }, { "text": "as throwing an error or not producing an output [167]. Current deepfake detection models often rely on a single cue from the generation process as the basis for detection, making them vulnerable to specific adversarial samples. Furthermore, relatively little work has focused on adversarial sample testing in terms of model generalization capability and detection evaluation.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 817 }, { "text": "2.5 Applications of Deepfake Detection Deepfake detection is essential for ensuring the integrity and security of digital media as synthetic content becomes increasingly realistic. Several major institutions have deployed detection tools in practical settings. Youtube incorporates automated deepfake classifiers into its content moderation", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 818 }, { "text": "pipeline to identify manipulated media before large-scale dissemination [168]. Financial institutions such as the Bank of China [169] employ deepfake-oriented liveness detection to prevent video- based impersonation fraud. These real-world deployments highlight the essential role of deepfake detection in protecting individuals, financial systems, and public information integrity.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 819 }, { "text": "8 Pei and Zhang, et al. Table 2. Overview of representative face swapping methods. Notations: \u278aSelf-build, \u278bCelebA-HQ, \u278cFFHQ, \u278dVGGFace2, \u278eVGGFace, \u278fCelebV, \u2790CelebA, \u2791VoxCeleb2, \u2792LFW, \u2793KoDF. Abbreviations: SIG- GRAPH (SIG.), EUROGRAPHICS (EG.), GANs (G.), VAEs (V.), Diffusion (D.), Split-up and Integration (SI.).", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 820 }, { "text": "Method Venue Dataset Categorize Limitation Highlight Traditional Graphics Blanz et al. [173] EG.\u201904 \u278a 3DMM Manual intervention, unnatural output. Early face-swapping efforts simplified manual interaction steps. Bitouk et al. [174] SIG.\u201908 \u278a SI. Manual intervention, attribute loss. A three-phase implementation framework with the help of a pre-constructed face database to match", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 821 }, { "text": "faces that are similar to the source face in terms of posture and lighting. Sunkavalli et al. [175] TOG\u201910 \u278a SI. Poor generalizability, frequent artifacts. Early work on face exchange was realized using image processing methods such as smooth histogram matching technique. Dale et al. [176] SIG.\u201911 \u278a", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 822 }, { "text": "3DMM Poor generalizability and output quality. Early work on face exchange, proposing an improved Poisson mixing approach to achieve face swapping in video through frame-by-frame face replacement. Lin et al. [177] ICME\u201912 [178] 3DMM Poor generalizability, frequent artifacts. An attempt to construct a personalized 3D head model to solve the artifact problem occurring in face", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 823 }, { "text": "swapping in large poses. Mosaddegh et al. [179] ACCV\u201914 [180][181] SI. Poor generalizability and output quality. A diverse form of face swapping where facial components can be targeted for replacement. Nirkin et al. [182] FG\u201918 [183] SI. Poor generalization ability and resolution. Transfer of expressions and poses by building some 3D variable models and training facial segmentation", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 824 }, { "text": "networks to maintain target facial occlusion. Generative Adversarial Network IPGAN [184] CVPR\u201918 [185] G.+V. Poor output image quality, frequent artifacts. Using two encoders to encode facial identity and attribute information separately for facial information decoupling and swapping. RSGAN [112] SIG.\u201918", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 825 }, { "text": "\u2790 G.+V. Loss of lighting information. Using two independent VAE modules to represent the latent spaces of the face and hair regions, respecti -vely, with the replacement of identity information in the latent space implemented. Sun et al. [113] ECCV\u201918 [186] G.+3DMM Poor ability to preserve face feature attributes.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 826 }, { "text": "Implementing in two stages: the first stage involves replacing the identity information of the face reg -ion, while the second stage achieves complete facial rendering. FSGAN [187] ICCV\u201919 [188] G. Poor ability to preserve face feature attributes. Two novel loss functions are introduced to refine the stitching in the face fusion phase following the", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 827 }, { "text": "swapping process. FaceShifter [189] CVPR\u201920 \u278a\u278b\u278c\u278e G. Poor ability to preserve face feature attributes. Face swapping is realized in two stages, the firstly AEI-Net improves the output image quality level, and the second HEAR-Net is targeted to focus on abnormal regions for image recovery. Zhu et al. [190]", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 828 }, { "text": "AAAI\u201920 \u278a G.+V. Inability to process facial contour information. First show of the applicability of deepfake to keypoint invariant de-identification work. SimSwap [191] MM\u201920 \u278a\u278d G.+V. Poor ability to preserve face feature attributes. ID modules and weak feature matching loss functions are proposed to find a balance", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 829 }, { "text": "between identity information replacement and attribute information retention. MegaFS [192] CVPR\u201921 \u278a\u278b\u278c G. Poor ability to preserve face feature attributes. The first method allows for face swapping on images with a resolution of one million pixels. HifiFace [193] IJCAI\u201921 \u278d G.+3DMM Uses a large number of parameters.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 830 }, { "text": "A 3D shape-aware identity extractor is proposed to achieve better retention of attribute information such as facial shape. FSGANv2 [194] TPAMI\u201922 \u278a[188] G. Unable to process posture differences effectively. An extension of the FSGAN method that combines Poisson optimization with perceptual loss enhances", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 831 }, { "text": "the output image facial details. FSLSD [116] CVPR\u201922 \u278a\u278b G. Poor ability to preserve face feature attributes. Potential semantic de-entanglement is realized to obtain facial structural attributes and appearance attributes in a hierarchical manner. Kim et al. [195] CVPR\u201922 \u278c\u278d G. Unable to process posture differences effectively. An identity embedder is proposed to enhance the training speed under supervision.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 832 }, { "text": "3DSwap [196] CVPR\u201923 \u278b\u278c G.+3DMM Unable to process posture differences effectively. A 3d-aware approach to the face-swapping task, de-entangling identity and attribute features in latent space to achieve identity replacement and attribute feature retention. BlendFace [13] ICCV\u201923 \u278a\u278c\u278d\u278f G. Unable to handle occlusion and extreme lighting. The identity features obtained from the de-entanglement are fed to the generator as an identity loss", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 833 }, { "text": "function, which guides the generator to generate an image to fit the source image identity information. FlowFace [197] AAAI\u201923 \u278a\u278b\u278c\u278dG.+3DMM Altered target image lighting details. It consists of face reshaping network and face exchange network, which better solves the influence of the difference between source and target face contours on the face exchange work.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 834 }, { "text": "S2Swap [198] MM\u201923 \u278a\u278b\u278c\u2791 G.+3D Poor ability to preserve face feature attributes. Achieving high-fidelity face swapping through semantic disentanglement and structural enhancement. StableSwap [199] TMM\u201924 \u278a\u278c G.+3D Unable to handle extreme skin color differences. Utilizing a multi-stage identity injection mechanism effectively combines facial features from both the", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 835 }, { "text": "source and target to produce high-fidelity face swapping. Difussion DiffSwap [200] CVPR\u201923 \u278a\u278c D. Poor ability to handle facial occlusion. Reenvisioning face swapping as conditional inpainting to harness the power of the diffusion model. Liu et al. [9] CVPR\u201924 \u278a\u278b\u278c D. Poor ability to preserve face feature attributes.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 836 }, { "text": "Conditional diffusion model introduces identity and expression encoders components, achieving a balan- ce between identity replacement and attribute preservation during the generation process. DiffFace [201] PR\u201925 \u278a\u278c D. Facial lighting attributes are altered. Claims to be the first diffusion model-based face exchange framework.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 837 }, { "text": "Baliah et al. [202] WACV\u201925 \u278c\u2790 D. Unable to handle extreme pose and expressions. Achieve improvements in identity fidelity, pose consistency, and model generalization capability. PixSwap [203] WACV\u201925 \u278b\u278c D. Unable to handle extreme pose and eyes control. Achieve accurately reflecting the source identity and generating more high-quality images.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 838 }, { "text": "Alternative Cui et al. [204] CVPR\u201923 \u278a\u278b Other Altered target image lighting details. Introducing a multiscale transformer network focusing on high-quality semantically aware corresponden- ces between source and target faces. TransFS [205] FG\u201923 \u278a\u278b\u2793 Other Unable to process posture differences effectively. The identity generator is designed to reconstruct high-resolution images of specific identities, and an", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 839 }, { "text": "attention mechanism is utilized to enhance the retention of identity information. Wang et al. [206] TMM\u201924 \u278a\u278b Other Poor ability to handle facial occlusion. A Global Residual Attribute-Preserving Encoder (GRAPE) is proposed, and a network flow considering the facial landmarks of the target face was introduced, achieving high-quality face swapping.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 840 }, { "text": "CanonSwap [207] ICCV\u201925 \u278a\u278e other Poor ability to preserve face feature attributes. Disentangles motion and appearance for video face swapping, enabling precise identity transfer with pre- served temporal dynamics . 2.6 Ethical and Societal Considerations The rapid development of deepfake technologies has raised major ethical and societal concerns,", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 841 }, { "text": "particularly regarding privacy invasion, identity misuse, and the spread of manipulated content. In response, several countries and regions have introduced regulatory frameworks for synthetic media. The European Union incorporates transparency requirements into the AI Act [170] and the Digital Services Act [171], mandating clear labeling of manipulated content and stricter platform", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 842 }, { "text": "obligations. China\u2019s \u201cProvisions on the Administration of Deep Synthesis Internet Information Services\u201d establish comprehensive rules on watermarking, traceability, identity verification, and platform accountability [172]. These policies illustrate a global movement toward responsible governance of synthetic media, emphasizing transparency and protection against misuse.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 843 }, { "text": "3 Deepfake Inspections: A Survey 3.1 Deepfake Generation 3.1.1 Face Swapping . In this section, we review face swapping methods from the perspective of basic architecture, which can be mainly divided into four categories and summarized in Tab. 2. \u2022 Traditional Graphics. As representative early implementations, traditional graphics methods", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 844 }, { "text": "Deepfake Generation and Detection: A Benchmark and Survey 9 fusion. Methods [174, 175, 182] ground in critical information matching and fusion are geared towards substituting corresponding regions by aligning key points within facial regions of interest (ROIs), such as the mouth, eyes, nose, and mouth, between the source and target images. Following", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 845 }, { "text": "this, additional procedures such as boundary blending and lighting adjustments are executed to produce the resulting image. Bitouk et al. [174] accomplish automated face replacement by constructing a substantial face database to locate faces with akin poses and lighting conditions for substitution. Meanwhile, Nirkin et al. [182] enhance keypoint matching and segmentation accuracy", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 846 }, { "text": "by incorporating a Fully Convolutional Network (FCN) into their method. 2) The construction of a 3D prior model for facial parameterization. Methods [173, 176] based on constructing a 3D prior and introducing a facial parameter model often involve building a facial parameter model using 3DMM technology based on a pre-collected face database. After matching the facial information of", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 847 }, { "text": "the source image with the constructed face model, specific modifications are made to the relevant parameters of the facial parameter model to generate a completely new face. Dale et al. [176] utilize 3DMM to track facial expressions in two videos, enabling face swapping in videos. Some methods [177, 208] explore scenarios involving significant pose differences between the source", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 848 }, { "text": "and target images. Lin et al. [177] construct a 3D face model from frontal faces, renderable in any pose. Guo et al. [208] utilize plane parameterization and affine transformation to establish a one-to-one dense mapping between 2D graphics. Traditional computer graphics methods solve basic face-swapping problems, exploring full automation to enhance generalization. However, these", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 849 }, { "text": "methods are constrained by the need for similarities in pose and lighting between source and target images. They also face challenges like low image resolution, modification of target attributes, and poor performance in extreme lighting and occlusion scenarios. \u2022 Generative Adversarial Network. GAN-based methods can be classified into seven categories:", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 850 }, { "text": "1) Early GAN-based methods [46, 190, 209, 210] address issues related to the similarity of pose and lighting between source and target images. DepthNets [209] combines GANs with 3DMM to map the source face to any target geometry, not limited to the geometric shape of the target template. This allows it to be less affected by differences in pose between the source and target", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 851 }, { "text": "faces. However, they face challenges in generalizing the trained model to unknown faces. 2) Improved Generalizability. To improve the model\u2019s generalization, many efforts [113, 187, 191] are made to explore solutions. Combining GANs with VAEs, the model [112, 184] encodes and processes different facial regions separately. FSGAN [187] integrates face reenactment with face", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 852 }, { "text": "swap, designing a facial blending network to mix two faces seamlessly. SimSwap [191] introduces an identity injection module to avoid integrating identity information into the decoder. These methods remain limited by low resolution, attribute degradation, and poor handling of facial occlusions. 3) Resolution Upgrading. Some methods [192, 211, 212] aim to improve the resolution of generated", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 853 }, { "text": "faces. MegaFS [192] introduces the first single-lens face swapping method at the million-pixel level. The encoder no longer compresses facial information but represents it in layers, achieving more detailed preservation. StyleIPSB [212] constrains semantic attribute codes within the subspace of StyleGAN, thereby fixing semantic information during face swapping to preserve pore-level details.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 854 }, { "text": "4) Geometric Detail Preservation. To capture and reproduce more facial geometric details, some methods [193, 197, 213, 214] introduce 3DMM into GANs, enabling the incorporation of 3D priors. HifiFace [193] introduces a novel 3D shape-aware identity extractor, replacing traditional face recognition networks to generate identity vectors that include precise shape information. Flow-", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 855 }, { "text": "Face [197] introduces a two-stage framework based on semantic guidance to achieve shape-aware face swapping. FlowFace++ [214] improves upon FlowFace by utilizing a pre-trained Mask Autoen- coder to convert face images into a fine-grained representation space shared between the target and source faces. It further enhances feature fusion for both source and target by introducing a", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 856 }, { "text": "10 Pei and Zhang, et al. cross-attention fusion module. However, most of the aforementioned methods often struggle to effectively handle occlusion issues. 5) Facial Masking Artifacts. Some methods [187, 189, 215, 216] have partially alleviated the arti- facts caused by facial occlusion. FSGAN [187] designs a restoration network to estimate missing", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 857 }, { "text": "pixels. E4S [215] redefines the face-swapping problem as a mask-exchanging problem for specific information. It utilizes a mask-guided injection module to perform face swapping in the latent space of StyleGAN. However, overall, the methods above have not thoroughly addressed the issue of artifacts in generated images under extreme occlusion conditions.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 858 }, { "text": "6) Trade-offs between Identity-replacement and Attribute-retention. In addition to the occlusion issues that need further handling, researchers [13, 217] discover that the balance between identity replacement and attribute preservation in generated images seems akin to a seesaw. Many meth- ods [195, 198, 218] explore the equilibrium between identity replacement and attribute retention.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 859 }, { "text": "StyleSwap [218] introduces a novel swapping guidance strategy, the ID reversal, to enhance the similarity of facial identity in the output. Shiohara et al. [13] propose BlendFace, using an identity encoder that extracts identity features from the source image and uses it as identity distance loss, guiding the generator to produce facial exchange results.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 860 }, { "text": "7) Model Light-weighting is also an important topic with profound implications for the widespread application of models. FastSwap [219] achieves this by innovating a decoder block called Triple Adaptive Normalization (TAN), effectively integrating identity information from the source image and pose information from the target image. XimSwap [12] modifies the design of convolutional", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 861 }, { "text": "blocks and the identity injection mechanism, successfully deploying on STM32H743. \u2022 Diffusion-based. The latest studies [9, 109, 200, 201] in this area produce promising generation results. DiffSwaps [200] redefines the face swapping problem as a conditional inpainting task. Liu et al. [9] introduce a multi-modal face generation framework and achieved this by introducing", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 862 }, { "text": "components such as balanced identity and expression encoders to the conditional diffusion model, striking a balance between identity replacement and attribute preservation during the generation process. As a novel facial generalist model, FaceX [109] can achieve various facial tasks, including face swapping and editing. Leveraging the pre-trained StableDiffusion [7] has significantly improved", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 863 }, { "text": "the quality and model training speed. Baliah et al. [202] achieves improved identity fidelity, pose consistency, and generalization through self-supervised inpainting, DDIM multi-step sampling, CLIP-based feature disentanglement, and mask shuffling. \u2022 Alternative Techniques. Some methods stand independently from the above classifications that", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 864 }, { "text": "are discussed here collectively. Fast Face-swap [220] views the identity swap task as a style transfer task, achieving its goals based on VGG-Net. However, this method has poor generalization. Some methods [204, 205] apply the Transformer architecture to face swapping tasks. Leveraging a facial encoder based on the Swin Transformer [53], TransFS [205] obtains rich facial features, enabling", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 865 }, { "text": "facial swapping in high-resolution images. CanonSwap [207] introduces a motion\u2013appearance dis- entangled video face swapping framework that enables accurate identity transfer while preserving fine-grained temporal dynamics. 3.1.2 Face Reenactment . This section reviews current methods from four points: 3DMM-based, landmark matching, face", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 866 }, { "text": "feature decoupling, and self-supervised learning. We summarize them in Tab. 3. \u2022 3DMM-based. Some methods [224, 225] utilize 3DMM to construct a facial parameter model as an intermediary for transferring information between the source and target. In particular, Face2Face\ud835\udf0c[224], based on 3DMM, consists of a u-shaped rendering network driven by head pose", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 867 }, { "text": "Deepfake Generation and Detection: A Benchmark and Survey 11 Table 3. Overview of face reenactment methods. Notations: \u278aVoxceleb, \u278bSelf-build, \u278cVoxceleb2, \u278dTalkingHead-1KH, \u278eCelebV-HQ, \u278fVFHQ, \u2790RaFD, \u2791VGGFace, \u2792CelebV, \u2793FFHQ. Expression (Exp). Method Venue Controllable object Dataset Highlight Based on 3DMM", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 868 }, { "text": "Kim et al. [221] TOG\u201918 Exp, Pose, Blink \u278b Using synthesized rendering images of a parameterized face model as input, creating lifelike video frames for the target actor. Kim et al. [222] TOG\u201919 Lip,Exp, Pose \u278b Built on a recurrent generative adversarial network, it employs a hierarchical neural face renderer to synthesize realistic video frames.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 869 }, { "text": "HeadGAN [223] ECCV\u201921 Exp, Pose \u278a Using 3DMM for facial modeling provides a 3D prior to the GAN, effectively guiding the generator to accurately recover pose and expression from the target frame. Face2Face\ud835\udf0c[224] ECCV\u201922 Exp, Pose \u278a Decoupling the actor\u2019s facial appearance and motion information with two separate encodings allows the network to learn facial app", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 870 }, { "text": "-earance and motion priors. PECHead [225] CVPR\u201923 Lip, Exp, Pose \u278c\u278d\u278e\u278f A novel multi-scale feature alignment module for motion perception is proposed to minimize distortion during motion transmission. Maskrenderer [226] PR\u201925 Lip, Exp, Pose \u278a The framework is robust to occlusion and large mismatches between Source and Driving facial structures.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 871 }, { "text": "Based on Landmark Matching Zakharov et al. [227] ICCV\u201919 Exp, Pose \u278a\u278c Proposed a meta-learning framework for adversarial generative models, reducing the required training data size. FReeNet [228] CVPR\u201920 Exp, Pose \u2790[68] A new triple perceptual loss is proposed to richly reproduce facial details of the face.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 872 }, { "text": "DG [14] CVPR\u201922 Exp, Pose \u278a\u278c\u2790 A proposed dual generator model network for large pose face reproduction. Doukas et al. [229] TPAMI\u201923 Exp, Pose, Gaze \u278a[230][231] Eye gaze control in the generated video is implemented to further enhance visual realism. MetaPortrait [232] CVPR\u201923 Exp, Pose \u278c By establishing dense facial keypoint matching, accurate deformation field prediction is achieved, and the model training is expedited", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 873 }, { "text": "based on the meta-learning philosophy. Yang et al. [233] AAAI\u201924 Exp, Pose \u278a\u278c\u278d The facial tri-plane is represented by canonical tri-plane, identity deformation, and motion components, achieving face reenactment without the need for 3D parameter model priors. FSRT[234] CVPR\u201924 Exp, Pose \u278a The Transformer-based encoder-decoder effectively encodes attributes and improves action transmission quality.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 874 }, { "text": "DiffusionAct[235] FG\u201925 Exp, Pose \u278a The nethod allows one-shot, self, and cross-subject reenactment, without requiring subject-specific fine-tuning. Based on Face Feature Decopling HiDe-NeRF [236] CVPR\u201923 Lip, Exp, Pose \u278a\u278c\u278d High-fidelity and free-viewing talking head synthesis using deformable neural radiation fields.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 875 }, { "text": "HyperReenact [15] ICCV\u201923 Exp, Pose \u278a\u278c Exploiting the effectiveness of hypernetworks in real image inversion tasks and extending them to real image manipulation. Stylemask [237] FG\u201923 Exp, Pose \u2793 This work optimizes a Mask Network and combines it with StyleGAN2\u2019s style potential space S in order to achieve the separation of", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 876 }, { "text": "facial pose and expression of the target image from the identity features of the source image. Bounareli et al. [238] IJCV\u201924 Exp, Pose \u278a\u278c In GANs\u2019 latent space, head pose and expression changes are decoupled, achieving near-real outputs through real image embedding. Chang et al. [239] AAAI\u201925 Exp, Pose", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 877 }, { "text": "\u278e[240] Introduce a StyleGAN2-based end-to-end framework for high-fidelity one-shot video reenactment at 1024 resolution, with conditional disentanglement and feature-space refinement improving accuracy and fine-detail preservation Based on Self-supervised ICface [241] WACV\u201920 Exp, Pose \u278a The model is decoupled and driven by interpretable control signals that can be obtained from multiple sources such as external driving", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 878 }, { "text": "videos and manual controls. Oorloff et al. [242] ICCV\u201923 Lip, Exp, Pose \u278e Identity and attribute decomposition are realized in StyleGAN2\u2019s latent space, and a cyclic manifold adjustment technique enhances facial reconstruction results. Xue et al. [243] TOMM\u201923 Exp, Pose \u278a\u278c High-fidelity facial generation is achieved by using information-rich Projected Normalized Coordinate Code (PNCC) and eye maps,", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 879 }, { "text": "replacing sparse facial landmark representations. such as incomplete attribute decoupling in facial reproduction tasks, PECHead [225] models facial expressions and pose movements. It combines self-supervised learning of landmarks with 3D facial landmarks and introduces a new motion-aware multi-scale feature alignment module to eliminate", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 880 }, { "text": "artifacts that may arise from facial motion. \u2022 Landmark Matching. This kind of methods [14, 229, 245\u2013247] aim to establish a mapping relationship between semantic objects in the facial regions of the driving source and the target source through landmarks. Based on this mapping relationship [227, 228], the transfer of facial", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 881 }, { "text": "movement information is achieved. To address the challenge of reproducing large head poses in facial reenactment, Xu et al. [14] propose a dual-generator network incorporating a 3D landmark detector into the model. Free-headgan [229] comprises a 3D keypoint estimator, an eye gaze estimator, and a generator built on the HeadGAN architecture. The 3D keypoint estimator addresses the regression", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 882 }, { "text": "of deformations related to 3D poses and expressions. The eye gaze estimator controls eye movement in videos, providing finer details. MetaPortrait [232] achieves accurate distortion field prediction through dense facial keypoint matching and accelerates model training based on meta-learning principles, delivering excellent results on limited datasets.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 883 }, { "text": "\u2022 Feature Decoupling. The latent feature decoupling and driving methods [15, 233, 236\u2013238] aims to disentangle facial features in the latent space of the driving video, replacing or mapping the corresponding latent information to achieve high-fidelity facial reproduction under specific condi- tions. HyperReenact [15] uses attribute decoupling, employing a hyper-network to refine source", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 884 }, { "text": "identity features and modify facial poses. StyleMask [237] separates facial pose and expression from the identity information of the source image by learning masks and blending corresponding channels in the pre-trained style space S of StyleGAN2. HiDe-NeRF [236] employs a deformable neural radiance field to represent a 3D scene, with a lightweight deformation module explicitly", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 885 }, { "text": "decoupling facial pose and expression attributes. \u2022 Self-supervised Learning. Self-supervised learning employs supervisory signals inferred from the intrinsic structure of the data, reducing the reliance on external data labels [241, 242, 278, 279]. Oorloff et al. [242] employs self-supervised methods to train an encoder, disentangling identity", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 886 }, { "text": "12 Pei and Zhang, et al. Table 4. Overview of representative talking face generation methods. Notations: \u278aLRW, \u278bVoxCeleb2, \u278cMEAD, \u278dSelf-build, \u278eLRS2, \u278fHDTF, \u2790LRS3, \u2791CREMA-D, \u2792VoxCeleb, \u2793FFHQ. Method Venue Dataset Limitation Highlight Audio / Text - Driven Chen et al. [100] ECCV\u201918 \u278a[248][249] Poor resolution, inability to control pose and emotional.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 887 }, { "text": "Proposed a novel generator network and a comprehensive model with four complementary losses, as well as a new audio-visual related loss function to guide video generation. Zhou et al. [250] AAAI\u201919 \u278a Inability to control pose and emotional variations. Generate high-quality talking face videos by disentangling audio-visual representations.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 888 }, { "text": "Chen et al. [115] CVPR\u201919 \u278a Inability to control pose and emotional variations. Proposed a cascaded approach, using facial landmarks as an intermediate high-level representation. Wav2Lip [101] ICMR\u201920 \u278a\u278e\u2790 Poor resolution, inability to control pose and emotional. A new evaluation framework and a dataset for training mouth synchronization are proposed.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 889 }, { "text": "MakeItTalk [251] TOG\u201920 \u278b Uable to control pose and emotional variations well. Separating content information and identity information from audio signals, combining LSTM and self -attention mechanism to enhance head movement coherence. Ji et al. [252] CVPR\u201921 \u278a\u278c Inability to control pose and emotional variations.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 890 }, { "text": "By breaking down the input audio sample into content and emotion embeddings, cross-reconstruction of emotional disentanglement creates facial landmarks with nuanced emotional content. SPACE [253] ICCV\u201923 \u278b\u278c Lack emotional and other latent attributes control Constructed a novel facial intermediate representation, achieving control overhead pose, blinking, and", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 891 }, { "text": "gaze direction. SadTalker [254] CVPR\u201923 \u278f\u2792 Lack emotional and other latent attributes control. Based on the idea of 3DMM and conditional VAE, 3D coefficients controlling facial motion and expre -ssion are generated from audio to realize the reproduction of accurate faces from audio. EmoTalk [255] ICCV\u201923", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 892 }, { "text": "\u278d\u278f Poor real-time performance and expression details. An emotion-entangled encoder and emotion-guided decoder enable emotion injection, with outputs generated using Blendshape and FLAME model rendering. TalkLip [256] CVPR\u201923 \u278a\u278e Inability to control pose and emotional variations. Pre-trained lip-reading experts are employed to penalize incorrect lip-reading predictions in the synth", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 893 }, { "text": "-esized videos. DR2 [17] WACV\u201924 \u278d Lack emotional and other latent attributes control. The model explored effective strategies for reducing the training workload. RADIO [257] WACV\u201924 \u278a\u278b\u278f Lack emotional and other latent attributes control. Introducing StyleGAN2 style modulation to adapt to human identity and utilizes ViT blocks to focus", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 894 }, { "text": "on facial attributes in the reference image. FT2TF [258] WACV\u201925 \u278e\u2790 Lack controllable emotional intensity regulation. A one-stage pipeline that generates realistic talking faces by integrating visual and textual input. Talkclip [259] TMM\u201925 \u278b\u278c\u278f Insufficient control over the intensity of emotional output. The method can use text to modulate expression intensity and edit expressions.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 895 }, { "text": "Multimodal PC-AVS [260] CVPR\u201921 \u278a\u278b Lack emotional and other latent attributes control. Introduction of pose-source video drive compensation to generate head motion in video. GC-AVT [261] CVPR\u201922 \u278b\u278c Poor resolution, unable to handl complex backgrounds. In addition to the source image, a gesture source, an expression source, and audio are introduced to", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 896 }, { "text": "drive the talking head generation. Yu et al. [262] TMM\u201922 \u278d Lack emotional and other latent attributes control. Fusion of audio and text inputs for more accurate lip movement and chin posture prediction. Xu et al. [263] CVPR\u201923 \u278c Insufficient control over the intensity of emotional output. Embedding textual, visual, and auditory emotional modalities into a unified space.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 897 }, { "text": "LipFormer [264] CVPR\u201923 \u278e\u2793 Poor ability to preserve face feature attributes. Propose retaining high-quality facial details obtained from pre-training in a codebook format and repro -ducing them by driving the encoded mapping relationship between audio and lip movements. Wang et al. [265] TPAMI\u201924 \u278b\u278c\u278f", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 898 }, { "text": "Unable to delicately control emotions. Using 3DMM as an intermediate variable to convey facial expressions and head movements, and intro -ducing additional reference videos to extract the desired speaking style. Diffusion DAE-Talker [266] MM\u201923 \u278d High model complexity. It replaces traditional manually crafted intermediate representations with data-driven latent representa", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 899 }, { "text": "-tions obtained from a DAE. Yu et al. [267] ICCV\u201923 \u278b\u2792 Poor resolution, high model complexity. Building a corresponding mapping between audio and non-lip representations and training using the diffusion model. Stypu\u0142kowski [268] WACV\u201924 \u278a\u2791 High model complexity, short video generation duration. The model incorporates motion frame and audio embedding information to capture past movements", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 900 }, { "text": "and future expressions, with an emphasis on the mouth region through an additional lip sync loss. EmoTalker [269] ICASSP\u201924 \u278c\u2791 High model complexity. It achieves emotion-editable talking face generation based on a conditional diffusion model. VASA-1 [270] NIPS\u201924 \u2792 High model complexity. Expressive and well-decoupled facial latent space has been constructed, and highly controllable, high", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 901 }, { "text": "-quality generation effects have been achieved based on the Diffusion Transformer. EmotiveTalk [271] CVPR\u201925 \u278c\u278f High model complexity. Enhances long-duration talking-face generation by introducing a visual-guided audio\u2013expression disen -tanglement strategy and an expressive diffusion backbone, enabling controllable emotional expression.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 902 }, { "text": "3D-Model AD-NeRF [272] ICCV\u201921 \u278d Inadequate control of emotions and latent attributes. The NeRF based approach achieves accurate reproduction of detailed facial components and generates the upper body region. DFRF [273] ECCV\u201922 \u278d Lack of emotional and other latent attributes control. Combining audio with 3D perceptual features and proposing an facial deformation module.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 903 }, { "text": "AE-NeRF [274] AAAI\u201924 \u278f Lack emotional and other latent attributes control. Facial modeling is divided into NeRF related to audio and unrelated to audio to enhance audio-visual lip synchronization and facial detail. SyncTalk [275] CVPR\u201924 \u278d Lack controllable emotional intensity regulation. The facial sync controller boosts component coordination, and a portrait generator corrects artifacts,", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 904 }, { "text": "enhancing video details. Ye et al. [276] ICLR\u201924 \u278b[74] Lack emotional control and occasional artifacts. Facial and audio information is separately represented using tri-plane, followed by rendering. The gen -erated results are further optimized based on the super-resolution network. Tang et al. [277]", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 905 }, { "text": "IJCV\u201925 [272][273] Lack controllable emotional intensity regulation. This framework enhances reenactment fidelity by decomposing portrait representations into low dimen -sional feature grids, enabling coherent audio-driven head motion and efficient torso modeling. pre-trained StyleGAN2. Zhang et al. [279] utilizes 3DMM to provide geometric guidance, employs", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 906 }, { "text": "pre-computed optical flow to guide motion field estimation, and relies on pre-computed occlusion maps to guide the perception and repair of occluded areas. 3.1.3 Talking Face Generation . In this section, we review current methods from three perspectives: audio/text driven, multimodal conditioned, diffusion-based, and 3D-model Technologies. We also summarize them in Tab. 4.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 907 }, { "text": "\u2022 Audio/Text Driven. Early methods [100, 114] perform poorly in terms of generalization and training complexity. After training, the models struggled to generalize to new individuals, requiring extensive conversational data for training new characters. Researchers [101, 115] propose their solutions from various perspectives. However, Most of these methods prioritize generating lip", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 908 }, { "text": "movements aligned with semantic information, overlooking essential aspects like identity and style, such as head pose changes and movement control, which are crucial in natural videos. To address this, MakeItTalk [251] decouples input audio information by predicting facial landmarks based on audio and obtaining semantic details on facial expressions and poses from audio signals.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 909 }, { "text": "SadTalker [254] extracts 3D motion coefficients for constructing a 3DMM from audio and uses this to modulate a new 3D perceptual facial rendering for generating head poses in talking videos. Additionally, some methods [92, 257, 280, 281] propose their improvement methods, and these will not be detailed one by one. In addition, the emotional expression varies for different texts", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 910 }, { "text": "Deepfake Generation and Detection: A Benchmark and Survey 13 between the driving information and corresponding emotions. EMMN [284] establishes an organic relationship between emotions and lip movements by extracting emotion embeddings from the audio signal, synthesizing overall facial expressions in talking faces rather than focusing solely on", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 911 }, { "text": "audio for facial expression synthesis. AMIGO [285] employs a sequence-to-sequence cross-modal emotion landmark generation network to generate vivid landmarks aided by audio information, ensuring that lips and emotions in the output image sequence are synchronized with the input audio. However, existing methods still lack effective control over the intensity of emotions. In addition,", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 912 }, { "text": "TalkCLIP [259] introduces style parameters, expanding the style categories for text-guided talking video generation. Zhong et al. [286] propose a two-stage framework, incorporating appearance priors during the generation process to enhance the model\u2019s ability to preserve attributes of the target face. DR2 [17] explores practical strategies for reducing the training workload.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 913 }, { "text": "\u2022 Multimodal Conditioned. To generate more realistic talking videos, some methods [260, 261, 263, 264] introduce additional modal information on top of audio-driven methods to guide facial pose and expression. GC-AVT [261] generates realistic talking videos by independently controlling head pose, audio information, and facial expressions. This approach introduces an expression source", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 914 }, { "text": "video, providing emotional information during the speech and the pose source video. However, the video quality falls below expectations, and it struggles to handle complex background changes. Xu et al. [263] integrate text, image, and audio-emotional modalities into a unified space to complement emotional content in textual information. Multimodal approaches have significantly enhanced", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 915 }, { "text": "the vividness of generated videos, but there is still room for exploration of organically combining information driven by different sources and modalities. \u2022 Diffusion-based. Recently, some methods [16, 266, 269] apply the Diffusion model to the task of talking face generation. For fine-grained talking video generation, DAE-Talker [266] replaces", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 916 }, { "text": "manually crafted intermediate representations, such as facial landmarks and 3DMM coefficients, with data-driven latent representations obtained from a Diffusion Autoencoder (DAE). The image decoder generates video frames based on predicted latent variables. EmoTalker [269] utilizes a conditional diffusion model for emotion-editable talking face generation. It introduces emotion", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 917 }, { "text": "intensity blocks and the FED dataset to enhance the model\u2019s understanding of complex emotions. Very recently, diffusion models are gaining prominence in talking face generation tasks [268, 287] and video generation tasks [143, 288]. Emo [287] directly predicts video from audio without the need for intermediate 3D components, achieving excellent results. However, the lack of explicit", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 918 }, { "text": "control signals may easily lead to unnecessary artifacts. Based on the Diffusion Transformer architecture, VASA-1 [270] finely encodes and reconstructs facial details, constructing an expressive and well-decoupled facial latent space. \u2022 3D-model Technologies. 3D models, exemplified by NeRF, are gaining traction in talking face", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 919 }, { "text": "generation [272\u2013274]. AD-NeRF [272] directly feeds features from the input audio signal into a conditional implicit function to generate a dynamic NeRF. AE-NeRF [274] employs a dual NeRF framework to separately model audio-related regions and audio-independent regions. Furthermore, some methods [275, 276] adopt Tri-Plane to represent facial and audio attributes. Synctalk [275]", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 920 }, { "text": "models and renders head motion using a tri-plane hash representation, and then further enhances the output quality using a portrait synchronization generator. Very recently, 3D Gaussian Splatting [289] also been widely applied to this task. Some method [312, 313] introduce 3DGS to achieve more refined facial reconstruction and motion details, aiming to address the issue of insufficient pose", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 921 }, { "text": "14 Pei and Zhang, et al. Table 5. Overview of facial attribute editing methods. Notations: \u278aFFHQ, \u278bCelebA, \u278cCelebA-HQ, \u278dCelebAMask-HQ, \u278eVoxCeleb, \u278fCelebAText-HQ, \u2790LFW, \u2791MM CelebA-HQ, \u2792CARLA, \u2793Multi-PIE. In addition, abbreviations are used in the table: SIGGRAPH (SIG.), GANs (G.), Diffusion (D.), Transformer (T.).", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 922 }, { "text": "Method Venue Categorize Dataset Highlight GeneGAN [290] BMVC\u201917 G. \u278b\u2793 In early attribute editing, separate models were trained for specific attributes. The key idea was to reassemble attribute vectors in the latent space, achieving successful editing. SC-FEGAN [291] ICCV\u201919 G. \u278c Users can generate high-quality edited output images by freely sketching parts of the source image.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 923 }, { "text": "AttGAN [292] TIP\u201919 G. \u278b Applying attribute classification constraints to generated images has validated the drawbacks of enforcing stringent attribute independence constraints in latent representations. Shen et al. [293] CVPR\u201920 G. \u278b Thoroughly investigated how to encode different semantics in the latent space and explored the disentanglement between various", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 924 }, { "text": "semantics to achieve precise control over facial attributes. Yao et al. [294] ICCV\u201921 G.+T. \u278c By integrating explicit decoupling terms and identity-consistent terms into the loss function, the preservation of facial identity information is improved, resulting in high-quality face editing in videos.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 925 }, { "text": "HifaFace [295] CVPR\u201921 G. \u278a Proposed a solution based on wavelet transform to address the issue of partial loss of attribute information when generating edited results due to \"cyclic consistency\" problems. Preechakul et al. [296] CVPR\u201922 D. \u278a When encoding images, it is divided into semantically meaningful parts and parts that represent the details of the image.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 926 }, { "text": "FENeRF [297] CVPR\u201922 G.+NeRF \u278a\u278d The introduction of semantic masks into the conditional radiance field enables finer image textures. GuidedStyle [298] NN\u201922 G. \u278a Generating faces after face editing is guided based on facial attribute classification. The introduction of a sparse attention mecha -nism enhances the manipulation of individual attribute styles.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 927 }, { "text": "FDNeRF [27] SIG.\u201922 G.+NeRF \u278e The introduction of the Conditional Feature Warping (CFW) module addresses the issue of temporal inconsistency caused by dynamic information in the process of face editing in videos. AnyFace [18] CVPR\u201922 G. \u278f\u2791 Proposed a dual-branch framework for text-driven facial editing, with coordination achieved between the two branches through", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 928 }, { "text": "a Cross-Modal Distillation (CMD) module. TransEditor [299] CVPR\u201922 G. \u278c\u278a Emphasizing dual-space GAN interaction\u2019s importance, a transformer architecture is introduced for improved interaction. Huang et al. [300] CVPR\u201923 D. \u278d Proposed the concept of assisted diffusion, integrating individual multimodalities to explore the complementarity between differe", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 929 }, { "text": "-nt modalities. Ozkan et al. [301] ICCV\u201923 G. \u278a The entangled attribute space is decomposed into conceptual and hierarchical latent spaces, and transformer network encoders are employed to modify information in the latent space. CIPS-3D++ [302] TPAMI\u201923 G.+NeRF \u278a\u2792 Replaced the convolutional architecture with an MLP (Multi-Layer Perceptron) architecture to achieve faster rendering speeds.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 930 }, { "text": "ClipFace [303] SIG.\u201923 G.+3DMM \u278a Learned texture generation from large-scale datasets, enhancing generator performance through generative adversarial training. TG-3DFace [304] ICCV\u201923 G. \u278c\u278f For different scenarios, two sets of text-to-face cross-modal alignment methods were designed with specific focuses.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 931 }, { "text": "VecGAN++ [305] TPAMI\u201923 G. \u278c Orthogonal constraint and disentanglement loss are used to decouple attribute vectors in the latent space. DiffusionRig [306] CVPR\u201923 D. \u278a 3DMM and diffusion model integration propose a two-stage method for learning personalized facial details. Kim et al. [307] CVPR\u201923 D.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 932 }, { "text": "\u278e Proposed a method for facial editing in videos based on the diffusion model. SDGAN et al. [308] AAAI\u201924 G. \u278c SDGAN introduces a semantic separation generator and a semantic mask alignment strategy, achieving satisfactory preservation of irrelevant details and precise attribute manipulation. FaceDNeRF [309]", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 933 }, { "text": "NIPS\u201924 D.+NeRF \u278a Creating and editing facial NeRFs with single-view images, text prompts, and target lighting. NeRFFaceEditing [310] TPAMI\u201925 G.+NeRF \u278a Disentangling geometry and appearance within a tri-plane NeRF through statistical tri-plane features and 3D semantic masks. M-LMPF [311] PR\u201925 G. \u278a", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 934 }, { "text": "M-LMPF achieving controllable attribute modification with strong privacy protection and superior editing fidelity. \u2022 Comprehensive Editing. Early facial attribute editing models [290, 314] often achieve editing for a single attribute through data-driven training. For instance, Shen et al. [314] propose learning", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 935 }, { "text": "the difference between pre-/post-operation images, represented as residual images, to achieve attribute-specific operations. However, single-attribute editing falls short of meeting expectations, and compression steps in the process often lead to a significant loss of image resolution, a common issue in early methods. The fundamental challenge in comprehensive editing and unrelated attribute", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 936 }, { "text": "modification is achieving complete attribute disentanglement. Many approaches [293\u2013295, 299] have explored this. E.g., HifaFace [295] identifies cycle consistency issues as the cause of facial attribute information loss that proposes a wavelet-based method for high-fidelity face editing, while TransEditor [299] introduces a dual-space GAN structure based on the transformer framework that", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 937 }, { "text": "improves image quality and attribute editing flexibility. \u2022 Irrelevant-attribute Retained. Another critical aspect of face editing is retaining as much target image information as possible in the generated images [19, 294, 298]. GuidedStyle [298] leverages attention mechanisms in StyleGAN [5] for the adaptive selection of style modifications", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 938 }, { "text": "for different image layers. IA-FaceS [19] embeds the face image to be edited into two branches of the model, where one branch calculates high-dimensional component-invariant content embed- ding to capture facial details, and the other branch provides low-dimensional component-specific embedding for component operations. Additionally, some approaches [26, 27, 297, 302] combine", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 939 }, { "text": "GANs with NeRF [24] for enhanced spatial awareness capabilities. Specifically, FENeRF [297] uses two decoupled latent codes to generate corresponding facial semantics and textures in a 3D volume with spatial alignment sharing the same geometry. CIPS-3D++ [302] enhances the model\u2019s training efficiency with a NeRF-based shallow 3D shape encoder and an MLP-based deep 2D image decoder.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 940 }, { "text": "Deepfake Generation and Detection: A Benchmark and Survey 15 Table 6. Overview of representative forgery detection methods. Notations: \u2780FF++, \u2781DFDC, \u2782Celeb-DF, \u2783Deeperforensics, \u2784Self-build, \u2785UADFV, \u2786Celeb-HQ, \u2787DFDCp, \u2788FFHQ, \u2789DFD. Method Venue Train Test Highlight Space Domain Gram-Net [316] CVPR\u201920", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 941 }, { "text": "\u2786\u2788 \u2786\u2788 The method posits that genuine faces and fake faces exhibit inconsistencies in texture details. Face X-ray [317] CVPR\u201920 \u2780 \u2780\u2781\u2782\u2789 Focusing on boundary artifacts of face fusion for forgery detection. Zhao et al. [318] CVPR\u201921 \u2780 \u2780\u2781\u2782 A texture enhancement module, an attention generation module, and a bilinear attention pooling mod", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 942 }, { "text": "-ule are proposed to focus on texture details. Nirkin et al. [319] TPAMI\u201921 \u2780 \u2780\u2781\u2782 Detecting swapped faces by comparing the facial region with its context (non-facial area). SBIs [320] CVPR\u201922 \u2780 \u2781\u2782\u2787\u2789 The belief that the more difficult to detect forged faces typically contain more generalized traces of forg", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 943 }, { "text": "-ery can encourage the model to learn a feature representation with greater generalization ability. LGrad [105] CVPR\u201923 \u2784 \u2784 The gradient is utilized to present generalized artifacts that are fed into the classifier to determine the truth of the image. NoiseDF [321] AAAI\u201923 \u2780 \u2780\u2781\u2782\u2783 Extracting noise traces and features from cropped faces and background squares in video frames.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 944 }, { "text": "Ba et al. [322] AAAI\u201924 \u2780\u2781\u2782 \u2780\u2781\u2782 Multiple non-overlapping local representations are extracted from the image for forgery detection. A local information loss function, based on information bottleneck theory, is proposed for constraint. UNITE [323] CVPR\u201925 \u2780\u2782\u2783\u2785 \u2780\u2782\u2783\u2785 Leveraging domain-agnostic features, attention-diversity regularization, and mixed-domain training,", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 945 }, { "text": "achieving great performance on both facial and fully synthetic manipulations. Time Domain FTCN [324] ICCV\u201921 \u2780 \u2780\u2781\u2782\u2783[189] It is believed that most face video forgeries are generated frame by frame. As each altered face is inde -pendently generated, this inevitably leads to noticeable flickering and discontinuity.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 946 }, { "text": "LipForensics [325] CVPR\u201921 \u2780 \u2780\u2781\u2782 Concern about temporal inconsistency of mouth movements in videos. M2TR [103] ICMR\u201922 \u2780 \u2780\u2781\u2782\u2789 Capturing local inconsistencies at different scales for forgery detection using a multiscale transformer. Gu et al. [326] AAAI\u201922 \u2780 \u2780\u2781\u2782[87] By densely sampling adjacent frames to pay attention to the inter-frame image inconsistency.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 947 }, { "text": "Yang et al. [33] TIFS\u201923 \u2780\u2781\u2782 \u2780\u2781\u2782 Treating detection as a graph classification problem and focusing on the relationship between the local image features across different frames. AVoiD-DF [38] TIFS\u201923 \u2781\u2784[85] \u2781\u2784[85] Multimodal forgery detection using audiovisual inconsistency. Choi et al. [327] CVPR\u201924", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 948 }, { "text": "\u2780\u2782\u2783 \u2780\u2782 Focus on the inconsistency of the style latent vectors between frames. Xu et al. [328] IJCV\u201924 \u2780\u2781\u2782\u2783 \u2780\u2781\u2782\u2783 Forgery detection is conducted by converting video clips into thumbnails containing both spatial and temporal information. Peng et al. [329] TIFS\u201924 \u2780\u2782\u2787 \u2780\u2782\u2787 Focuse on inter-frame gaze angles, extracting gaze informations and employing spatio-temporal feature", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 949 }, { "text": "aggregation to combine temporal, spatial, and texture features for detection and classification. Frequency FDFL [36] CVPR\u201921 \u2780 \u2780 Propose an adaptive frequency feature generation module to extract differential features from different frequency bands in a learnable manner. HFI-Net [330] TIFS\u201922 \u2780 \u2781\u2782\u2783\u2785[79] Notice that the forgery flaws used to distinguish between real and fake faces are concentrated in the", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 950 }, { "text": "mid- and high-frequency spectrum. Guo et al. [331] TIFS\u201923 \u2780\u2781 \u2780\u2781\u2782 Designing a backbone network for Deepfake detection with space-frequency interaction convolution. Tan et al. [332] AAAI\u201924 \u2784 \u2780\u2784 A lightweight frequency-domain learning network is proposed to constrain classifier operation within the frequency domain.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 951 }, { "text": "WMamba [333] MM\u201925 \u2780 \u2781\u2782\u2787 Introduces a Mamba-based wavelet feature extractor that leverages dynamic contour convolution and efficient long-range modeling to capture fine-grained, globally distributed forgery cues. WaveDIF [334] CVPR\u201925 \u2780\u2782 \u2780\u2782 A lightweight frequency-domain detector that leverages DFT-based denoising and wavelet subband en", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 952 }, { "text": "-ergy analysis to achieve competitive accuracy on both intra- and cross-dataset evaluations. Data Driven Dang et al. [335] CVPR\u201920 \u2784 \u2782\u2785 Utilizing attention mechanisms to handle the feature maps of the detection model. Zhao et al. [336] ICCV\u201921 \u2780 \u2780\u2781\u2782\u2783\u2787\u2789Proposes pairwise self-consistent learning for training CNN to extract these source features and detect", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 953 }, { "text": "deep vacation images. Finfer [337] AAAI\u201922 \u2780 \u2780\u2782\u2787[87] Based on an autoregressive model, using the facial representation of the current frame to predict the facial representation of future frames. Huang et al. [104] CVPR\u201923 \u2780 \u2780\u2781\u2782\u2789[189] A new implicit identity-driven face exchange detection framework is proposed.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 954 }, { "text": "HiFi-Net [338] CVPR\u201923 \u2784 \u2784 Converting forgery detection and localization into a hierarchical fine-grained classification problem. Zhai et al. [339] ICCV\u201923 [340] [341][342] Weakly supervised image processing detection is proposed such that only binary image level labels (real or tampered) are required for training.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 955 }, { "text": "introduces two text-to-face cross-modal alignment techniques, including global contrastive learning and fine-grained alignment modules, to enhance the high semantic consistency. \u2022 Diffusion-based. Diffusion-based models have been introduced into facial attribute editing [296, 300, 306, 307] and achieve excellent results. Huang et al. [300] propose a collaborative diffusion", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 956 }, { "text": "framework, utilizing multiple pre-trained unimodal diffusion models together for multimodal face generation and editing. DiffusionRig [306] conditions the initial 3D face model, which helps preserve facial identity information during personalized editing of facial appearance. 3.2 Forgery Detection In this section, we review current forgery detection techniques based on the type of detection cues,", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 957 }, { "text": "categorizing them into three: Space Domain, Time Domain, Frequency Domain, and Data-Driven. We also summarize the detailed information about popular methods in Tab. 6. 3.2.1 Space Domain . \u2022 Image-level Inconsistency. The generation process of forged images often involves partial alterations rather than global generation, leading to common local differences in non-globally", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 958 }, { "text": "16 Pei and Zhang, et al. as criteria for determining whether an image is forged, such as color [30], saturation [343], arti- facts [318, 320, 344], gradient variations [105], etc. Specifically, RECCE [344] considers shadow generation from a training perspective, utilizing the learned representations on actual samples to", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 959 }, { "text": "identify image reconstruction differences. LGrad [105] utilizes a pre-trained transformation model, converting images to gradients to visualize general artifacts and subsequently classifying based on these representations. In addition, some works focus on detection based on differences in facial and", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 960 }, { "text": "non-facial regions [319], as well as the fine-grained details of image textures [316, 345]. Recently, Ba et al. [322] focuse not only on the discordance in a single image region but also on the detection of fused local representation information from multiple non-overlapping areas. \u2022Local Noise Inconsistency. Image forgery may involve adding, modifying, or removing content", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 961 }, { "text": "in the image, potentially altering the noise distribution in the image. Detection methods based on noise aim to identify such local or even global differences in the image. Zhou et al. [31] propose a dual-stream structure, combining GoogleNet with a triplet network to focus on tampering artifacts and local noise in images. NoiseDF [321] specializes in identifying underlying noise traces left", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 962 }, { "text": "behind in Deepfake videos, introducing an efficient and novel Multi-Head Relative Interaction with depth-wise separable convolutions to enhance detection performance. 3.2.2 Time Domain . \u2022 Abnormal Physiological Information. Forgery videos often overlook the authentic physiolog- ical features of humans, failing to achieve overall consistency with authentic individuals. Therefore,", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 963 }, { "text": "some methods focus on assessing the plausibility of the physiological features of the generated faces in videos. Li et al. [78] detect blinking and blink frequency in videos as criteria for determining the video\u2019s authenticity. Yang et al. [89] focuses on the inconsistency of head poses in videos, comparing the differences between head poses estimated using all facial landmarks and those", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 964 }, { "text": "estimated using only the landmarks in the central region. Peng et al. [329] focuse on inter-frame gaze angles, obtaining gaze characteristics of each video frame and using a spatio-temporal feature aggregator to combine temporal gaze features, spatial attribute features, and spatial texture features", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 965 }, { "text": "as the basis for detection and classification. \u2022 Inter-Frame Inconsistency. Methods [32, 33, 324, 326\u2013328] based on inter-frame inconsistency for forgery detection aim to uncover differences in images between adjacent frames or frames with specific temporal spans. Gu et al. [326] focuse on inter-frame image inconsistency by densely", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 966 }, { "text": "sampling adjacent frames, while Yin et al. [32] design a Dynamic Fine-grained Difference Cap- turing module and a Multi-Scale Spatio-Temporal Aggregation module to cooperatively model spatio-temporal inconsistencies. Yang et al. [33] approach forgery detection as a graph classifi- cation problem, emphasizing the relationship information between facial regions to capture the", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 967 }, { "text": "relationships among local features across different frames. Choi et al. [327] discover that the style variables in each frame of Deepfake work change. Based on this, they developed a style attention module to focus on the inconsistency of the style latent variables between frames. \u2022 Multimodal Inconsistency. The core idea behind multimodal detection algorithms is to make", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 968 }, { "text": "judgments based on the flow of prior information from multiple attributes rather than solely consid- ering the image or audio differences of individual characteristics in each frame. The consideration of audio-visual modal inconsistency has received extensive research in various methods [37, 38, 346].", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 969 }, { "text": "POI-Forensics [346] proposes a deep forgery detection method based on audio-visual authentication, utilizing contrastive learning to learn the most distinctive embeddings for each identity in moving facial and audio segments. AVoiD-DF [38] embeds spatiotemporal information in a spatiotempo- ral encoder and employs a multimodal joint decoder to fuse multimodal features and learn their", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 970 }, { "text": "Deepfake Generation and Detection: A Benchmark and Survey 17 detecting fake faces using static and dynamic auditory ear characteristics. Indeed, multimodal detection methods are currently a hotspot in forgery detection research. 3.2.3 Frequency Domain . Frequency domain-based forgery detection methods transform image time-domain information", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 971 }, { "text": "into the frequency domain. Works [35, 36, 330, 332, 348] utilize statistical measures of periodic features, frequency components, and frequency characteristic distributions, either globally or in local regions, as evaluation metrics for forgery detection. Specifically, F3-Net [35] proposes a dual- branch framework. One frequency-aware branch utilizes Frequency-aware Image Decomposition", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 972 }, { "text": "(FAD) to learn subtle forgery patterns in suspicious images. In contrast, the other branch aims to extract high-level semantics from Local Frequency Statistics (LFS) to describe the frequency-aware statistical differences between real and forged faces. HFI-Net [330] consists of a dual-branch network", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 973 }, { "text": "and four Global-Local Interaction (GLI) modules. It effectively explores multi-level frequency artifacts, obtaining frequency-related forgery clues for face detection. Tan et al. [332] introduce a novel frequency-aware approach called FreqNet, which focuses on the high-frequency information of images and combines it with a frequency-domain learning module to learn source-independent", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 974 }, { "text": "features. Furthermore, some approaches combine spatial, temporal, and frequency domains for joint consideration [331, 349], and Guo et al. [331] design a spatial-frequency interaction convolution to construct a novel backbone network for Deepfake detection. 3.2.4 Data Driven . Data-driven forgery detection [335, 337, 338, 350] focuses on learning specific patterns and features", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 975 }, { "text": "from extensive image or video datasets to distinguish between genuine and potentially manipulated images. Some methods [351, 352] believe that images generated by specific models possess unique model fingerprints. Based on this belief, forgery detection can be achieved by focusing on the model\u2019s training. In addition, FakeSpotter [352] introduces the Neuron Coverage Criterion to capture", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 976 }, { "text": "layer-wise neuron activation behavior. It monitors the neural behavior of a deep face recognition system through a binary classifier to detect fake faces. There are also methods [104, 336] that attempt to classify the sources of different components in an image. For instance, Huang et al. [104] think that the difference between explicit and implicit identity helps detect face swapping.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 977 }, { "text": "3.3 Specific Related Domains 3.3.1 Face Super-resolution . \u2022 Convolutional Neural Networks. Early works [118, 353, 354] on facial super-resolution based on CNNs aims to leverage the powerful representational capabilities of CNNs to learn the mapping relationship between low-resolution and high-resolution images from training samples. Depending", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 978 }, { "text": "on whether they focus on local details of the image, they can be divided into global methods [117, 118], local methods [353], and mixed methods [354]. \u2022 Generative Adversarial Network. GAN aims to achieve the optimal output result through an adversarial process between the generator and the discriminator. This type of method [355, 356]", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 979 }, { "text": "currently dominates the field for flexible and efficient architecture. 3.3.2 Portrait Style Transfer . \u2022 Generative Adversarial Network. The most mature style transfer algorithm is the GAN-based approach [152\u2013154]. However, due to the relatively poor stability of GANs, it is common for the generated images to contain artifacts and unreasonable components. 3DAvatarGAN [152] bridges", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 980 }, { "text": "the pre-trained 3D-GAN in the source domain with the 2D-GAN trained on an artistic dataset to achieve cross-domain generation. Scenimefy [154] utilizes semantic constraints provided by text models like CLIP to guide StyleGAN generation and applies patch-based contrastive style loss to enhance stylization and fine details further.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 981 }, { "text": "18 Pei and Zhang, et al. \u2022 Diffusion-based. Diffusion-based methods [155, 156, 357] represent the generative process of cross-domain image transfer using diffusion processes. DiffusionGAN3D [357] combines 3DGAN with a diffusion model from text to graphics, introducing relative distance loss and learnable", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 982 }, { "text": "tri-planes for specific scenarios to further enhance cross-domain transformation accuracy. 3.3.3 Body Animation . \u2022 Generative Adversarial Network. GAN-based approaches [141, 142] aim to train a model to generate images whose conditional distribution resembles the target domain, thus transferring information from reference images to target images. CASD [141] is based on a style distribution", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 983 }, { "text": "module using a cross-attention mechanism, facilitating pose transfer between source semantic styles and target poses. However, existing methods still rely considerably on training samples, and exhibit decreased performance when dealing with actions in rare poses. \u2022 Diffusion-based. The task of body animation using diffusion models aims to utilize diffusion", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 984 }, { "text": "processes to generate the propagation and interaction of movements between body parts based on a reference source. This approach [143, 144] represents a current hot topic in research and implementation. LEO [143] focuses on the spatiotemporal continuity between generated actions, employing the Latent Motion Diffusion Model to represent motion as a series of flow graphs during", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 985 }, { "text": "the generation process. 3.3.4 Makeup Transfer . \u2022 Graphics-based Approaches. Before the advent of neural networks, traditional computer graphics methods [358, 359] relied on image-gradient editing and physics-based operations to interpret makeup semantics. These approaches decomposed an input image into multiple layers", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 986 }, { "text": "and warped the reference makeup image onto the target using facial landmarks. However, the reliance on manually designed operators often led to unnatural results, including visible artifacts and unintended alterations to background regions. \u2022 Generative Adversarial Network. Early deep learning-based methods [160] aim at fully auto-", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 987 }, { "text": "matic makeup transfer. However, these methods [159, 360] exhibit poor performance when faced with significant differences in pose and expression between the source and target faces and are unable to handle extreme makeup scenarios well. Some methods [157, 161, 162] proposes their so- lutions, PSGAN++ [161] comprises the Makeup Distillation Network, Attentive Makeup Morphing", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 988 }, { "text": "module, Style Transfer Network, and Identity Extraction Network, further enhancing the ability of PSGAN [157] to perform targeted makeup transfer with detail preservation. ELeGANt [163], CUMTGAN [165], and HT-ASE [158] explore the preservation of detailed information. 4 Benchmark We introduce the evaluation metrics commonly used for each deepfake task, followed by a summary", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 989 }, { "text": "of the performance of representative methods on widely adopted datasets based on the results reported in their original papers. Given the variations in testing setups, and evaluation criteria across different approaches, we strive to present fair and consistent comparisons in each table. 4.1 Metrics", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 990 }, { "text": "\u2022 Face Swapping. The most commonly used objective evaluation metrics for face swapping include ID Ret, Expression Error, Pose Error, and FID. ID Ret is calculated by a pre-trained face recognition model [97], measuring the Euclidean distance between the generated face and the source face. A higher ID Ret indicates better preservation of identity information. Expression and pose errors", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 991 }, { "text": "quantify the differences in expression and pose between the generated face and the source face. These metrics are evaluated using a pose estimator [99] and a 3D facial model [361], extracting expression and pose vectors for the generated and source faces. Lower values for expression error and pose error indicate higher facial expression and pose similarity between the swapped face and", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 992 }, { "text": "Deepfake Generation and Detection: A Benchmark and Survey 19 the source face. FID [95] is used to assess image quality, with lower FID values indicating that the generated images closely resemble authentic facial images in appearance. \u2022 Face Reenactment. Face reenactment commonly uses consistent evaluation metrics, including", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 993 }, { "text": "CSIM, SSIM [93], PSNR, LPIPS [94], LMD [100], and FID. CSIM describes the cosine similarity between the generated and source faces, calculated by ArcFace [362], with higher values indicating better performance. SSIM, PSNR, LPIPS, and FID are used to measure the quality of synthesized images. SSIM measures the structural similarity between two images, with higher values indicating", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 994 }, { "text": "a closer resemblance to natural images. PSNR quantifies the ratio between a signal\u2019s maximum possible power and noise\u2019s power, indicating higher quality for higher values. LPIPS assesses reconstruction fidelity using a pre-trained AlexNet [48] to extract feature maps for similarity score computation. As mentioned earlier, FID is used to evaluate image quality. LMD assesses the accuracy", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 995 }, { "text": "of lip shape in generated images or videos, with lower values indicating better model performance. \u2022 Talking Face Generation. Expanding upon face reenactment metrics, talking face generation incorporates additional metrics, including M/F-LMD, Sync, LSE-C, and LSE-D. The LSE-C and LSE-D are usually used to measure lip synchronization effectiveness [101]. The landmark distances", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 996 }, { "text": "on the mouth (M-LMD) [115] and the confidence score of SyncNet (Sync) measure synchronization between the generated lip motion and the input audio. F-LMD computes the difference in the average distance of all landmarks between predictions and ground truth (GT) as a measure to assess the generated expression. In addition, there are some meaningful metrics, such as LSE-C and", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 997 }, { "text": "LSE-D for measuring lip synchronization effectiveness [101], AVD [236] for evaluating identity preservation performance, AUCON [247] for assessing facial pose and expression jointly, and AGD [229] for evaluating eye gaze changes. These newly proposed evaluation metrics enrich the performance assessment system by targeting various aspects of the model\u2019s performance.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 998 }, { "text": "\u2022 Facial Attribute Editing. The standard evaluation metrics used in face attribute manipulation are FID, LPIPS [94], KID [96], PSNR and SSIM. KID is one of the image quality assessment metrics commonly used in face editing work and other generative modeling tasks to quantify the difference in distribution between the generated image and the actual image, with lower KID values indicating", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 999 }, { "text": "better model performance. Some text-guide work will also use the CLIP Score to measure the consistency between the output image and the text, calculated as the cosine similarity between the normalized image and the text embedding. Higher values of CLIP Score indicate better consistency of the generated image with the corresponding text sentence.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1000 }, { "text": "4.2 Benchmark Protocol In the absence of a unified benchmark, this survey adopts commonly used datasets and evaluation metrics for each task to establish a reference benchmarking protocol. All reported results are drawn directly from the original publications; therefore, the benchmark is intended for performance", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1001 }, { "text": "presentation rather than strict comparative evaluation. \u2022 Face Swapping. This task is primarily evaluated on the FF++ [80] dataset. A test set is constructed by uniformly sampling 10 frames from each of 1,000 videos (10,000 images in total), and performance is measured using ID Retention, Expression Error, Pose Error, and FID. However, differences", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1002 }, { "text": "in training datasets may affect result comparability. In this survey, we follow this evaluation protocol and explicitly report experimental settings that may impact fairness. Tab. 7 summarizes the results reported in the original publications and documents training data differences. Notably, RAFSwap [11] and Xu et al. [116] adopted the MegaFS [192] preprocessing protocol for FF++, while", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1003 }, { "text": "SimSwap [191] and CanonSwap [207] applied pixel-threshold filtering to improve data quality. \u2022 Face Reenactment. This task is evaluated under three settings: self-reenactment, cross-identity reenactment, and quality assessment. VoxCeleb [69] is used for self- and cross-identity reenactment, while VoxCeleb2 [70] is adopted for quality assessment. For self-reenactment, the first frame of", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1004 }, { "text": "20 Pei and Zhang, et al. Table 7. Results of representative face swapping methods on FF++. Notations: \u278aCelebA-HQ, \u278bFFHQ, \u278cVGGFace, \u278dVGGFace2, \u278eVoxCeleb2. Methods Train Test: FF++ ID Ret.(%)\u2191 Exp Err.\u2193 Pose Err.\u2193 FID\u2193 FaceShifter [189] \u278a\u278b\u278c 97.38 2.06 2.96 - SimSwap [191] \u278d 92.83 - 1.53 - HifiFace [193]", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1005 }, { "text": "\u278d 98.48 - 2.63 - RAFSwap [11] \u278a 96.70 2.92 2.53 - Xu et al. [116] \u278b 90.05 2.79 2.46 - DiffSwap [200] \u278b 98.54 5.35 2.45 2.16 FlowFace [197] \u278a\u278b\u278d 99.26 - 2.66 - StyleIPSB [212] \u278b 95.05 2.23 3.58 - StyleSwap [218] \u278c\u278e 97.05 5.28 1.56 2.72 WSC-Swap [213] \u278a\u278b\u278c 99.88 5.01 1.51 - DiffFace [201] \u278b - 2.71 2.35", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1006 }, { "text": "- CanonSwap [207] \u278c 99.78 - 1.59 - Table 8. Results of representative face reenactment methods on VoxCeleb for the self-reenactment. No- tations: \u278aVoxCeleb, \u278bVoxCeleb2, \u278cETH-Xgaze, \u278dGaze360 [230], \u278eMPIIGaze [231], \u278fTalkingHead- 1KH. In addition, we use gray to represent data that is partially uncertain.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1007 }, { "text": "Methods Train Test: VoxCeleb CSIM\u2191 PSNR\u2191 LPIPIS\u2193 FID\u2193 SSIM\u2191 HyperReenact [15] \u278a 0.710 - 0.230 27.10 - DG [14] \u278a 0.831 - - 22.10 0.761 AVFR-GAN [363] \u278a - 32.20 - 8.48 0.824 Free-HeadGAN [229] \u278a\u278c\u278d\u278e 0.810 22.16 0.100 35.40 - HiDe-NeRF [236] \u278a\u278b\u2790 0.931 21.90 0.084 - 0.862 DiffusionAct [235] \u278a\u278b 0.690 19.70", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1008 }, { "text": "0.240 - 0.830 Table 9. Results of representative face reenactment methods on VoxCeleb for the cross-identity reen- actment. Notations: \u278aVoxCeleb, \u278bVoxCeleb2, \u278cETH-Xgaze, \u278dGaze360 [230], \u278eMPIIGaze [231], \u278fTalkingHead-1KH. Methods Train Test : VoxCeleb CSIM\u2191 AVD\u2193 AUCON\u2191 FID\u2193 AGD\u2193 HyperReenact [15] \u278a 0.680", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1009 }, { "text": "- - - - AVFR-GAN [363] \u278a - - - 9.05 - Free-HeadGAN [229] \u278a\u278c\u278d\u278e 0.789 - - 53.90 13.1 HiDe-NeRF [236] \u278a\u278b\u278f 0.786 0.012 0.971 57.00 - DiffusionAct [235] \u278a\u278b 0.600 - - - - Table 10. Results of representative face reenactment methods on VoxCeleb2 for quality assessment. Nota- tions: \u278aVoxCeleb, \u278bVoxCeleb2, \u278cLRW, \u278dCelebV-HQ,", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1010 }, { "text": "\u278eTalkingHead-1KH. Methods Train Test: VoxCeleb2 CSIM\u2191 PSNR\u2191 LMD\u2193 FID\u2193 SSIM\u2191 PC-AVS [260] \u278b\u278c - - 6.880 - 0.886 GC-AVT [261] \u278b - - 2.757 - 0.739 Wang et al. [92] \u278b - 28.92 1.830 - 0.830 PECHead [225] \u278b\u278d\u278e 1.590 - 23.05 - DG [14] \u278a 0.721 - - 51.79 0.540 HiDe-NeRF [236] \u278a\u278e 0.787 - - 61.00 - each VoxCeleb test video serves as the source image and the remaining frames as driving images.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1011 }, { "text": "For cross-identity reenactment, 35 video pairs with different identities are randomly selected from the VoxCeleb test set. Self-reenactment is evaluated using CSIM, PSNR, LPIPS, FID, and SSIM; cross-identity reenactment uses CSIM, AVD, AUCON, FID, and AGD; and quality assessment adopts CSIM, PSNR, LMD, FID, and SSIM. Tab. 8\u201310 report the original results and document differences", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1012 }, { "text": "in training datasets. \u2022 Talking Face Generation. Since the MEAD [71] dataset provides explicit emotion annotations, it is commonly adopted as the benchmark for this task. As shown in Tab. 11, this survey follows this protocol and uses MEAD as the benchmark dataset. The evaluation metrics include CSIM,", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1013 }, { "text": "LMD, M/F-LMD, Sync, FID, PSNR, and SSIM. We report the results from the original publications of the selected methods and document their corresponding training dataset configurations. It is worth noting that the selected methods differ in their use of the MEAD dataset, and the validation set splits are not standardized across studies.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1014 }, { "text": "\u2022 Facial Attribute Editing. This task is typically evaluated in terms of facial reconstruction capability and FID, yet lacks a unified evaluation protocol. Therefore, as shown in Tab. 12 and Tab. 13, this survey does not impose fixed training or validation splits; instead, it reports representative", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1015 }, { "text": "methods\u2019 PSNR, LPIPS, SSIM, and FID results as presented in the original publications. \u2022 Forgery Detection. This task requires evaluation under both self-dataset and cross-dataset settings, with AUC and ACC as the commonly used metrics. In this survey, self-dataset perfor- mance is evaluated on the FF++ dataset, with results separately reported for FF++ (HQ) and FF++", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1016 }, { "text": "Deepfake Generation and Detection: A Benchmark and Survey 21 Table 11. Evaluation results of the models involved in talking face generation on the MEAD dataset. No- tations: \u278aMEAD, \u278bLRW, \u278cVoxCeleb2, \u278dHDTF. Method Train Test: MEAD CSIM\u2191 LMD\u2193 M/F-LMD\u2193 Sync\u2191 FID\u2193 PSNR/SSIM\u2191 Xu et al. [263] \u278a 0.83 2.36", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1017 }, { "text": "- 3.500 15.91 30.09/0.850 EMMN [284] \u278a\u278b - - 2.780/2.870 3.570 - 29.38/0.660 AMIGO [285] \u278a\u278c - 2.44 2.140/2.440 - 19.59 30.29/0.820 SLIGO [282] \u278a 0.88 1.83 - 3.690 - -/0.790 Gan et al. [283] \u278a\u278c - - 2.250/2.470 - 19.69 21.75/0.680 SPACE [253] \u278a\u278c - - - 3.610 11.68 - TalkCLIP [259] \u278a - - 3.601/2.415 3.773", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1018 }, { "text": "- -/0.829 MTHM [364] \u278a - - - - 22.38 21.38/0.660 Table 12. Results of facial reconstruction capabilities in representative facial attribute editing work. Nota- tions: \u278aFFHQ, \u278bVoxCeleb, \u278cVoxCeleb2, \u278dCelebA- HQ. Methods Type Train Test PSNR\u2191 LPIPS\u2193 SSIM\u2191 Konpat et al. [296] Difussion \u278a \u278d - 0.0110 0.991", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1019 }, { "text": "Kim et al. [307] Difussion \u278a \u278a - 0.0450 0.922 FDNeRF [27] GANs+NeRF \u278a \u278a - 0.1420 0.821 HairNeRF [365] GANs+NeRF \u278c \u278c 31.84 0.1060 0.827 IA-FaceS [19] GANs \u278a\u278d \u278d 22.34 0.2240 0.642 IA-FaceS [19] GANs \u278a\u278d \u278a 22.43 0.0384 0.659 r-FaceS [366] GANs \u278a\u278d \u278d 30.69 0.0220 - Table 13. FID evaluation of different methods. No-", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1020 }, { "text": "tations: \u278aFFHQ, \u278bCelebA-HQ, \u278cMM CelebA-HQ, \u278dCelebA, \u278eCelebAText-HQ, \u278fSelf-build. Methods Type Train Test FID\u2193 FENeRF [297] GANs+NeRF \u278a \u278b 12.10 FENeRF [297] GANs+NeRF \u278a \u278a 28.20 AnyFace [18] GANs \u278b \u278e 56.75 AnyFace [18] GANs \u278b \u278c 50.56 TextFace [315] GANs \u278b \u278b 22.81 TG-3DFace [304] GANs \u278c \u278e 52.21 TG-3DFace [304]", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1021 }, { "text": "GANs \u278c \u278c 39.02 HifaFace [295] GANs \u278a\u278b \u278a\u278b 4.04 GuidedStyle [298] GANs \u278d \u278f 41.79 NeRFFaceEditing [310] GANs \u278a \u278a 6.00 Table 14. Results of the self-dataset performance on FF++. HQ (Mild compression), LQ (Heavy compres- sion). Methods Train FF++ (LQ) FF++ (HQ) ACC(%) AUC(%) ACC(%) AUC(%) F3-Net [35] FF++", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1022 }, { "text": "93.02 95.80 98.95 99.30 Masi et al. [349] FF++ 86.34 - 96.43 - Zhao et al. [318] FF++ 88.69 90.40 97.60 99.29 FDFL [36] FF++ 89.00 92.40 96.69 99.30 LipForensics [325] FF++ 94.20 98.10 98.80 99.70 RECCE [344] FF++ 91.03 95.02 97.06 99.32 Guo et al. [350] FF++ 92.76 96.85 99.24 99.75 MRL [33] FF++ 91.81", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1023 }, { "text": "96.18 93.82 98.27 4.3 Main Results on Deepfake Generation \u2022 Results on Face Swapping. Tab. 7 displays the performance evaluation results of some repre- sentative models on the Face Swapping task using the FF++ [80] dataset. WSC-Swap [213] captures external facial attribute information and internal identity features through two independent en-", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1024 }, { "text": "coders, enabling strong identity preservation and stable facial pose retention. However, it shows sub-optimal performance on facial expression error metrics. In addition, the method has been reported to suffer from certain attribute loss during facial attribute transfer. \u2022 Results on Face Reenactment. Tab. 8 and Tab. 9 show the performance evaluation results", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1025 }, { "text": "on the VoxCeleb [69] dataset for self-reenactment and cross-subject reenactment, respectively. AVFR-GAN [363] achieves better performance by using the multimodal modeling. Tab. 10 presents quality assessment results on the VoxCeleb2 dataset. HiDe-NeRF [236] represents 3D scenes using canonical appearance fields and implicit deformation fields. It achieves accurate facial attribute", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1026 }, { "text": "modeling by explicitly decoupling facial pose and expression attributes using a deformation module. \u2022 Results on Talking Face Generation. Tab. 11 displays the performance results of various talking face generation approaches on the MEAD [71] dataset since 2023. AMIGO [285] achieves promising results, which utilizes seq2seq to generate facial landmarks for emotion tagging, matching mouth", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1027 }, { "text": "movements with emotional features. Additionally, it employs a landmark-to-image translation network to create facial images with fine textures. \u2022 Results on Facial Attribute Editing. Tab. 12 evaluates the quality level of generated images using FID, and Tab. 13 assesses facial reconstruction capabilities using PSNR, LPIPS, and SSIM. Due", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1028 }, { "text": "to different training and testing datasets, quantitative fair comparisons are not possible that just serve as performance demonstrations. 4.4 Main Results on Forgery Detection Tab. 14 presents the ACC and AUC metrics for some detection models trained on FF++ [80] and tested on FF++ (HQ) and FF++ (LQ). LipForensics [325] exhibits robust performance on the strongly", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1029 }, { "text": "22 Pei and Zhang, et al. Table 15. Results of cross-dataset performance evaluation on four datasets DFDC, Celeb-DF (CDF), Celeb- DFv2 (CDFv2), and DeeperForensics-1.0 (DFo). Evaluation indicator is AUC. Notations: \u278aFF++, \u278bFF++(Real), \u278cFF++(HQ), \u278dFF++(LQ), \u278eSelf-build, \u278fSR-DF [103], \u2790DFDCp, \u2791FakeAVCeleb, \u2792DefakeAVMiT.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1030 }, { "text": "Methods Train DFDC CDF CDFv2 DFo Face X-ray [317] \u278a\u278e 80.92 80.58 - - Zhao et al. [336] \u278b 67.52 98.30 90.03 99.41 Zheng et al. [324] \u278a 74.00 - 86.90 98.80 LipForensics [325] \u278a 73.50 - 82.40 97.60 M2TR [103] \u278f - 82.10 - - M2TR [103] \u278a - 68.20 - - RECCE [344] \u278a 69.06 68.71 - - SBIs [320] \u2790 - - 90.79 -", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1031 }, { "text": "SBIs [320] \u278a 72.42 - 93.18 - RealForensics [37] \u278a 75.90 - 86.90 99.30 Guo et al. [350] \u278c 81.65 84.97 - - Yin et al. [32] \u278d 73.08 71.36 - - MRL [33] \u278d 71.53 83.58 - - AVoiD-DF [38] \u2791 80.60 - - - WMamba [333] \u278b 82.97 96.29 - - compressed FF++ (LQ), while Guo et al. [350] perform best on FF++ (HQ). Tab. 15 shows the", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1032 }, { "text": "cross-dataset evaluation. AVoiD-DF [38] and Zhao et al. [336] demonstrate excellent generalization ability, but there is still significant room for improvement in these datasets. However, overall, there is room for improvement in the evaluation performance of forgery detection models on the DFDC. 5", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1033 }, { "text": "Future Prospects \u2022 Face Swapping. Generalization is a significant issue in face swapping models. While many models demonstrate excellent performance on their training sets, there is often noticeable performance degradation when applied to different datasets during testing. In addition, beyond the common", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1034 }, { "text": "evaluation metrics, various face swapping works employ different evaluation metric systems, lacking a unified evaluation protocol. This absence hinders researchers from intuitively assessing model performance. Therefore, establishing comprehensive experimental and evaluation frameworks is crucial for fair comparisons, and driving progress in the field.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1035 }, { "text": "\u2022 Face Reenactment. Existing methods have room for improvement, facing three main challenges: convenience, authenticity, and security. Many approaches struggle to balance lightweight deploy- ment and generating high-quality reenactment effects, hindering the widespread adoption of facial reenactment technology in industries. Moreover, several methods claim to achieve high-quality", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1036 }, { "text": "facial reenactment, but they exhibit visible degradation in output during rapid pose changes or ex- treme lighting conditions in driving videos. Additionally, the computational complexity, consuming significant time and system resources, poses substantial challenges for practical applications. \u2022 Talking Face Generation. Current methods aim to improve the realism of generated conver-", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1037 }, { "text": "sational videos, yet they still lack fine-grained control over emotional dynamics. The alignment between emotional intonation and audio\u2013semantic content remains imprecise, and the modulation of emotional intensity is overly coarse. Moreover, the realistic coupling between head pose and facial expression is often under-modeled. Finally, for text or audio conveying strong emotions,", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1038 }, { "text": "existing approaches still produce noticeable artifacts in head motion. \u2022 Facial Attribute Editing. Currently, mainstream facial attribute editing employs the decoupling concept based on GANs and diffusion models are gradually being introduced into this field. The primary challenge is effectively separating the facial attributes to prevent unintended processing of", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1039 }, { "text": "other facial features during attribute editing. Additionally, there needs to be a universally accepted benchmark dataset and evaluation framework for fair assessments of facial editing. \u2022 Forgery Detection. With the rapid development of facial forgery techniques, the central challenge in face forgery detection technology is accurately identifying various forgery methods using a", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1040 }, { "text": "Deepfake Generation and Detection: A Benchmark and Survey 23 single detection model. Simultaneously, ensuring that the model exhibits robustness when detecting forgeries in the presence of disturbances such as compression is crucial. Most detection models follow a generic approach targeting common operational steps of a specific forgery method, such", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1041 }, { "text": "as the integration phase in face swapping or assessing temporal inconsistencies, but this manner limits the model\u2019s generalization capabilities. Moreover, as forgery techniques evolve, forged videos may evade detection by introducing interference during the detection process. \u2022 Discussion. Future development of deepfake generation and detection technologies will move", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1042 }, { "text": "toward more robust, adaptive and multimodal systems. For generation, integrating visual, auditory and textual cues can strengthen semantic consistency and identity preservation, while larger and higher-quality datasets will support better cross-domain generalization. Reinforcement learning and feedback-driven optimization may further enhance visual fidelity and temporal coherence. Detection", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1043 }, { "text": "research is expected to benefit from multimodal evidence aggregation and finer temporal modeling, leveraging cues such as audio\u2013video synchronization, speech semantics and physiological signals. Self-supervised and weakly supervised learning will also help reduce dependence on extensive annotations as new manipulation types continue to emerge. Deepfake generation and detection", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1044 }, { "text": "will expand into applications including digital humans, telepresence, assistive media creation and privacy-preserving data synthesis. As adoption grows, ethical risks related to identity misuse, privacy violations and unauthorized content generation become more pressing. Incorporating transparency mechanisms such as watermarking and provenance metadata, along with strong", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1045 }, { "text": "misuse detection and clear governance guidelines, will be essential for ensuring the safe and responsible deployment of deepfake technologies. 6 Conclusion This survey comprehensively reviews the latest developments in the field of deepfake generation and detection, which is the first to cover a variety of related fields thoroughly and discusses the", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1046 }, { "text": "latest technologies such as diffusion. Specifically, this paper covers an overview of basic background knowledge, including concepts of research tasks, the development of generative models and neural networks, and other information from closely related fields. Subsequently, we summarize the technical approaches adopted by different methods in the mainstream four generation and one", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1047 }, { "text": "detection fields, and classify and discuss the methods from a technical perspective. In addition, we strive to fairly organize and benchmark the representative methods in each field. Finally, we summarize the current challenges and future research directions for each field. References [1] Tong Sha, Wei Zhang, Tong Shen, Zhoujun Li, and Tao Mei. Deep person generation: A survey from the perspective", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1048 }, { "text": "of face, pose, and cloth synthesis. CSUR, 2023. [2] Diederik P Kingma and Max Welling. Auto-encoding variational bayes. stat, 2014. [3] Kihyuk Sohn, Honglak Lee, and Xinchen Yan. Learning structured output representation using deep conditional generative models. In NeurIPS, 2015. [4] Aaron Van Den Oord, Oriol Vinyals, et al. Neural discrete representation learning. In NeurIPS, 2017.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1049 }, { "text": "Zion English, Vikram Voleti, Adam Letts, et al. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv, 2023. [8] Yuwei Guo, Ceyuan Yang, Anyi Rao, Yaohui Wang, Yu Qiao, Dahua Lin, and Bo Dai. Animatediff: Animate your personalized text-to-image diffusion models without specific tuning. In ICLR, 2024.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1050 }, { "text": "24 Pei and Zhang, et al. [10] Simone Barattin, Christos Tzelepis, Ioannis Patras, and Nicu Sebe. Attribute-preserving face dataset anonymization via latent code optimization. In CVPR, 2023. [11] Chao Xu, Jiangning Zhang, Miao Hua, Qian He, Zili Yi, and Yong Liu. Region-aware face swapping. In CVPR, 2022.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1051 }, { "text": "transformers. In ICAART, 2024. [17] Chenxu Zhang, Chao Wang, Yifan Zhao, Shuo Cheng, Linjie Luo, and Xiaohu Guo. Dr2: Disentangled recurrent representation learning for data-efficient speech video synthesis. In WACV, 2024. [18] Jianxin Sun, Qiyao Deng, Qi Li, Muyi Sun, Min Ren, and Zhenan Sun. Anyface: Free-style text-to-face synthesis and", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1052 }, { "text": "manipulation. In CVPR, 2022. [19] Wenjing Huang, Shikui Tu, and Lei Xu. Ia-faces: A bidirectional method for semantic face editing. NN, 2023. [20] Youxin Pang, Yong Zhang, Weize Quan, Yanbo Fan, Xiaodong Cun, Ying Shan, and Dong-ming Yan. Dpe: Disentan- glement of pose and expression for general video portrait editing. In CVPR, 2023.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1053 }, { "text": "synthesis with latent diffusion models. In CVPR, 2022. [24] B Mildenhall, PP Srinivasan, M Tancik, JT Barron, R Ramamoorthi, and R Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In ECCV, 2020. [25] Kyle Gao, Yina Gao, Hongjie He, Dening Lu, Linlin Xu, and Jonathan Li. Nerf: Neural radiance field in 3d vision, a", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1054 }, { "text": "comprehensive review. arXiv, 2022. [26] Kaiwen Jiang, Shu-Yu Chen, Feng-Lin Liu, Hongbo Fu, and Lin Gao. Nerffaceediting: Disentangled face editing in neural radiance fields. In SIGGRAPH, 2022. [27] Jingbo Zhang, Xiaoyu Li, Ziyu Wan, Can Wang, and Jing Liao. Fdnerf: Few-shot dynamic neural radiance fields for", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1055 }, { "text": "face reconstruction and expression editing. In SIGGRAPH, 2022. [28] Shuaibo Li, Shibiao Xu, Wei Ma, and Qiu Zong. Image manipulation localization using attentional cross-domain cnn features. TNNLS, 2021. [29] Shuaibo Li, Wei Ma, Jianwei Guo, Shibiao Xu, Benchong Li, and Xiaopeng Zhang. Unionformer: Unified-learning", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1056 }, { "text": "transformer with multi-view representation for image manipulation detection and localization. In CVPR, 2024. [30] Peisong He, Haoliang Li, and Hongxia Wang. Detection of fake images via the ensemble of deep representations from multi color spaces. In ICIP, 2019. [31] Peng Zhou, Xintong Han, Vlad I Morariu, and Larry S Davis. Two-stream neural networks for tampered face detection.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1057 }, { "text": "In CVPRW, 2017. [32] Qilin Yin, Wei Lu, Bin Li, and Jiwu Huang. Dynamic difference learning with spatio-temporal correlation for deepfake video detection. TIFS, 2023. [33] Ziming Yang, Jian Liang, Yuting Xu, Xiao-Yu Zhang, and Ran He. Masked relation learning for deepfake detection. TIFS, 2023. [34] Hafsa Ilyas, Ali Javed, and Khalid Mahmood Malik. Avfakenet: A unified end-to-end dense swin transformer deep", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1058 }, { "text": "learning model for audio\u2013visual deepfakes detection. Applied Soft Computing, 2023. [35] Yuyang Qian, Guojun Yin, Lu Sheng, Zixuan Chen, and Jing Shao. Thinking in frequency: Face forgery detection by mining frequency-aware clues. In ECCV, 2020. [36] Jiaming Li, Hongtao Xie, Jiahong Li, Zhongyuan Wang, and Yongdong Zhang. Frequency-aware discriminative", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1059 }, { "text": "feature learning supervised by single-center loss for face forgery detection. In CVPR, 2021. [37] Alexandros Haliassos, Rodrigo Mira, Stavros Petridis, and Maja Pantic. Leveraging real talking faces via self- supervision for robust forgery detection. In CVPR, 2022. [38] Wenyuan Yang, Xiaoyu Zhou, Zhikai Chen, Bofei Guo, Zhongjie Ba, Zhihua Xia, Xiaochun Cao, and Kui Ren. Avoid-df:", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1060 }, { "text": "Audio-visual joint learning for detecting deepfake. TIFS, 2023. [39] Yisroel Mirsky and Wenke Lee. The creation and detection of deepfakes: A survey. CSUR, 2021. [40] Kishan Vyas, Preksha Pareek, Ruchi Jayaswal, and Shruti Patil. Analysing the landscape of deep fake detection: A survey. IJISAE, 2024.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1061 }, { "text": "Deepfake Generation and Detection: A Benchmark and Survey 25 [41] Yunfan Liu, Qi Li, Qiyao Deng, Zhenan Sun, and Ming-Hsuan Yang. Gan-based facial attribute manipulation. TPAMI, 2023. [42] Andrew Melnik, Maksim Miasayedzenkau, Dzianis Makaravets, Dzianis Pirshtuk, Eren Akbulut, Dennis Holzmann, Tarek Renusch, Gustav Reichert, and Helge Ritter. Face generation and editing with stylegan: A survey. TPAMI, 2024.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1062 }, { "text": "networks. In ICCV, 2017. [46] Jianmin Bao, Dong Chen, Fang Wen, Houqiang Li, and Gang Hua. Cvae-gan: fine-grained image generation through asymmetric training. In ICCV, 2017. [47] Yann LeCun, L\u00e9on Bottou, Yoshua Bengio, and Patrick Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 1998.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1063 }, { "text": "2020s. In CVPR, 2022. [51] Jiangning Zhang, Xiangtai Li, Jian Li, Liang Liu, Zhucun Xue, Boshen Zhang, Zhengkai Jiang, Tianxin Huang, Yabiao Wang, and Chengjie Wang. Rethinking mobile block for efficient attention-based models. In ICCV, 2023. [52] Jiangning Zhang, Xiangtai Li, Yabiao Wang, Chengjie Wang, Yibo Yang, Yong Liu, and Dacheng Tao. Eatformer:", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1064 }, { "text": "Improving vision transformer inspired by evolutionary algorithm. IJCV, 2024. [53] Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In ICCV, 2021. [54] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, \u0141ukasz Kaiser, and Illia", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1065 }, { "text": "Polosukhin. Attention is all you need. In NeurIPS, 2017. [55] Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Trans- formers for image recognition at scale. In ICLR, 2021.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1066 }, { "text": "man, Fr\u00e9do Durand, Joshua B Tenenbaum, and Vincent Sitzmann. Neural groundplans: Persistent neural scene representations from a single image. In ICLR, 2023. [58] Muhammad Zubair Irshad, Sergey Zakharov, Katherine Liu, Vitor Guizilini, Thomas Kollar, Adrien Gaidon, Zsolt Kira, and Rares Ambrus. Neo 360: Neural fields for sparse view synthesis of outdoor scenes. In ICCV, 2023.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1067 }, { "text": "26 Pei and Zhang, et al. [70] Joon Son Chung, Arsha Nagrani, and Andrew Zisserman. Voxceleb2: Deep speaker recognition. In Interspecch, 2018. [71] Kaisiyuan Wang, Qianyi Wu, Linsen Song, Zhuoqian Yang, Wayne Wu, Chen Qian, Ran He, Yu Qiao, and Chen Change Loy. Mead: A large-scale audio-visual dataset for emotional talking-face generation. In ECCV, 2020.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1068 }, { "text": "A large-scale video facial attributes dataset. In ECCV, 2022. [75] Ting-Chun Wang, Arun Mallya, and Ming-Yu Liu. One-shot free-view neural talking-head synthesis for video conferencing. In CVPR, 2021. [76] Triantafyllos Afouras, Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman. Deep audio-visual", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1069 }, { "text": "speech recognition. TPAMI, 2018. [77] Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman. Lrs3-ted: a large-scale dataset for visual speech recognition. arXiv, 2018. [78] Yuezun Li, Ming-Ching Chang, and Siwei Lyu. In ictu oculi: Exposing ai generated fake face videos by detecting eye blinking. In WIFS, 2018.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1070 }, { "text": "dataset for deepfake detection. In ACM MM, 2020. [88] Patrick Kwon, Jaeseong You, Gyuhyeon Nam, Sungwoo Park, and Gyeongsu Chae. Kodf: A large-scale korean deepfake detection dataset. In ICCV, 2021. [89] Xin Yang, Yuezun Li, and Siwei Lyu. Exposing deep fakes using inconsistent head poses. In ICASSP, 2019.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1071 }, { "text": "Deepfake Generation and Detection: A Benchmark and Survey 27 [101] KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar. A lip sync expert is all you need for speech to lip generation in the wild. In ACM MM, 2020. [102] Binh M Le and Simon S Woo. Quality-agnostic deepfake detection with intra-model collaborative learning. In ICCV,", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1072 }, { "text": "2023. [103] Junke Wang, Zuxuan Wu, Wenhao Ouyang, Xintong Han, Jingjing Chen, Yu-Gang Jiang, and Ser-Nam Li. M2tr: Multi-modal multi-scale transformers for deepfake detection. In ICML, 2022. [104] Baojin Huang, Zhongyuan Wang, Jifan Yang, Jiaxin Ai, Qin Zou, Qian Wang, and Dengpan Ye. Implicit identity", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1073 }, { "text": "driven deepfake face swapping detection. In CVPR, 2023. [105] Chuangchuang Tan, Yao Zhao, Shikui Wei, Guanghua Gu, and Yunchao Wei. Learning on gradients: Generalized artifacts representation for gan-generated images detection. In CVPR, 2023. [106] Xinru Chen, Chengbo Dong, Jiaqi Ji, Juan Cao, and Xirong Li. Image manipulation detection by multi-view multi-scale", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1074 }, { "text": "supervision. In ICCV, 2021. [107] Justin Johnson, Alexandre Alahi, and Li Fei-Fei. Perceptual losses for real-time style transfer and super-resolution. In ECCV, 2016. [108] Changyong Shu, Hemao Wu, Hang Zhou, Jiaming Liu, Zhibin Hong, Changxing Ding, Junyu Han, Jingtuo Liu, Errui Ding, and Jingdong Wang. Few-shot head swapping in the wild. In CVPR, 2022.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1075 }, { "text": "learning simultaneously image colorization and super-resolution. In AAAI, 2022. [111] Brian B Moser, Federico Raue, Stanislav Frolov, Sebastian Palacio, J\u00f6rn Hees, and Andreas Dengel. Hitchhiker\u2019s guide to super-resolution: Introduction and recent advances. TPAMI, 2023. [112] Ryota Natsume and Yatagawa. Rsgan: face swapping and editing using face and hair representation in latent spaces.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1076 }, { "text": "In SIGGRAPH, 2018. [113] Qianru Sun, Ayush Tewari, Weipeng Xu, Mario Fritz, Christian Theobalt, and Bernt Schiele. A hybrid model for identity obfuscation by face replacement. In ECCV, 2018. [114] Bo Fan, Lijuan Wang, and Frank K Soong. Photo-real talking head with deep bidirectional lstm. In ICASSP, 2015.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1077 }, { "text": "Diffusion models, image super-resolution and everything: A survey. In AAAI, 2024. [122] Yukai Shi, Guanbin Li, Qingxing Cao, Keze Wang, and Liang Lin. Face hallucination by attentive sequence optimization with reinforcement learning. TPAMI, 2019. [123] Kui Jiang, Zhongyuan Wang, Peng Yi, Guangcheng Wang, Ke Gu, and Junjun Jiang. Atmfn: Adaptive-threshold-based", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1078 }, { "text": "multi-model fusion network for compressed face hallucination. TMM, 2019. [124] Araceli Morales, Gemma Piella, and Federico M Sukno. Survey on 3d face reconstruction from uncalibrated images. Computer Science Review, 2021. [125] Sahil Sharma and Vijay Kumar. 3d face reconstruction in deep learning era: A survey. ACME, 2022.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1079 }, { "text": "A model-free approach. In ECCV, 2018. [129] Shubham Tulsiani, Tinghui Zhou, Alexei A Efros, and Jitendra Malik. Multi-view supervision for single-view reconstruction via differentiable ray consistency. In CVPR, 2017. [130] Yifan Xing, Rahul Tewari, and Paulo Mendonca. A self-supervised bootstrap method for single-image 3d face", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1080 }, { "text": "28 Pei and Zhang, et al. [131] Ira Kemelmacher-Shlizerman and Ronen Basri. 3d face reconstruction from a single image using a single reference face shape. TPAMI, 2010. [132] Luo Jiang, Juyong Zhang, Bailin Deng, Hao Li, and Ligang Liu. 3d face reconstruction with geometry details from a single image. TIP, 2018.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1081 }, { "text": "Fernando De la Torre. Personalized face inpainting with diffusion models by parallel visual attention. In WACV, 2024. [139] Peiqing Yang, Shangchen Zhou, Qingyi Tao, and Chen Change Loy. Pgdiff: Guiding diffusion models for versatile face restoration via partial guidance. In NeurIPS, 2024. [140] Wing-Yin Yu, Lai-Man Po, Ray CC Cheung, Yuzhi Zhao, Yu Xue, and Kun Li. Bidirectionally deformable motion", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1082 }, { "text": "modulation for video-based human pose transfer. In ICCV, 2023. [141] Xinyue Zhou, Mingyu Yin, Xinyuan Chen, Li Sun, Changxin Gao, and Qingli Li. Cross attention based style distribution for controllable person image synthesis. In ECCV, 2022. [142] Jian Zhao and Hui Zhang. Thin-plate spline motion model for image animation. In CVPR, 2022.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1083 }, { "text": "Shou. Magicanimate: Temporally consistent human image animation using diffusion model. In CVPR, 2024. [147] Shuai Yang, Liming Jiang, Ziwei Liu, and Chen Change Loy. Vtoonify: Controllable high-resolution portrait video style transfer. ACM TOG, 2022. [148] Xu Peng, Junwei Zhu, Boyuan Jiang, Ying Tai, Donghao Luo, Jiangning Zhang, Wei Lin, Taisong Jin, Chengjie Wang,", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1084 }, { "text": "and Rongrong Ji. Portraitbooth: A versatile portrait model for fast identity-preserved personalization. In CVPR, 2024. [149] Qiang Cai, Mengxu Ma, Chen Wang, and Haisheng Li. Image neural style transfer: A review. Computers and Electrical Engineering, 2023. [150] Juan C P\u00e9rez, Thu Nguyen-Phuoc, Chen Cao, Artsiom Sanakoyeu, Tomas Simon, Pablo Arbel\u00e1ez, Bernard Ghanem,", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1085 }, { "text": "Ali Thabet, and Albert Pumarola. Styleavatar: Stylizing animatable head avatars. In WACV, 2024. [151] Jianwei Feng and Prateek Singhal. 3d face style transfer with a hybrid solution of nerf and mesh rasterization. In WACV, 2024. [152] Rameen Abdal, Hsin-Ying Lee, Peihao Zhu, Menglei Chai, Aliaksandr Siarohin, Peter Wonka, and Sergey Tulyakov.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1086 }, { "text": "3davatargan: Bridging domains for personalized editable avatars. In CVPR, 2023. [153] Shiyao Xu, Lingzhi Li, Li Shen, Yifang Men, and Zhouhui Lian. Your3demoji: Creating personalized emojis via one-shot 3d-aware cartoon avatar synthesis. In SIGGRAPH, 2022. [154] Yuxin Jiang, Liming Jiang, Shuai Yang, and Chen Change Loy. Scenimefy: Learning to craft anime scene via semi-", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1087 }, { "text": "supervised image-to-image translation. In ICCV, 2023. [155] Jin Liu, Huaibo Huang, Chao Jin, and Ran He. Portrait diffusion: Training-free face stylization with chain-of-painting. arXiv, 2023. [156] Jiwan Hur, Jaehyun Choi, Gyojin Han, Dong-Jae Lee, and Junmo Kim. Expanding expressiveness of diffusion models", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1088 }, { "text": "with limited data via self-distillation based fine-tuning. In WACV, 2024. [157] Wentao Jiang, Si Liu, Chen Gao, Jie Cao, Ran He, Jiashi Feng, and Shuicheng Yan. Psgan: Pose and expression robust spatial-aware gan for customizable makeup transfer. In CVPR, 2020. [158] Mingxiu Li, Wei Yu, Qinglin Liu, Zonglin Li, Ru Li, Bineng Zhong, and Shengping Zhang. Hybrid transformers with", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1089 }, { "text": "Deepfake Generation and Detection: A Benchmark and Survey 29 [160] Qiao Gu, Guanzhi Wang, Mang Tik Chiu, Yu-Wing Tai, and Chi-Keung Tang. Ladn: Local adversarial disentangling network for facial makeup and de-makeup. In ICCV, 2019. [161] Si Liu, Wentao Jiang, Chen Gao, Ran He, Jiashi Feng, Bo Li, and Shuicheng Yan. Psgan++: robust detail-preserving", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1090 }, { "text": "makeup transfer and removal. TPAMI, 2021. [162] Xiaojing Zhong, Xinyi Huang, Zhonghua Wu, Guosheng Lin, and Qingyao Wu. Sara: Controllable makeup transfer with spatial alignment and region-adaptive normalization. arXiv, 2023. [163] Chenyu Yang, Wanrong He, Yingqing Xu, and Yang Gao. Elegant: Exquisite and locally editable gan for makeup", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1091 }, { "text": "transfer. In ECCV, 2022. [164] Qixin Yan, Chunle Guo, Jixin Zhao, Yuekun Dai, Chen Change Loy, and Chongyi Li. Beautyrec: Robust, efficient, and component-specific makeup transfer. In CVPR, 2023. [165] Miao Hao, Guanghua Gu, Hao Fu, Chang Liu, and Dong Cui. Cumtgan: An instance-level controllable u-net gan for", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1092 }, { "text": "facial makeup transfer. Knowledge-Based Systems, 2022. [166] Sicong Han, Chenhao Lin, Chao Shen, Qian Wang, and Xiaohong Guan. Interpreting adversarial examples in deep learning: A review. CSUR, 2023. [167] Chaoran Yuan, Xiaobin Liu, and Zhengyuan Zhang. The current status and progress of adversarial examples attacks.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1093 }, { "text": "In CISCE, 2021. [168] https://blog.youtube/inside-youtube/. Accessed: Nov.19, 2025. [169] https://baijiahao.baidu.com/s?id=1810168520550712093&wfr=spider&for=pc. Accessed: Nov.19, 2025. [170] https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai. Accessed: Nov.19, 2025. [171] https://digital-strategy.ec.europa.eu/en/policies/digital-services-act-package. Accessed: Nov.19, 2025.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1094 }, { "text": "replacing faces in photographs. In SIGGRAPH, 2008. [175] Kalyan Sunkavalli, Micah K Johnson, Wojciech Matusik, and Hanspeter Pfister. Multi-scale image harmonization. ACM TOG, 2010. [176] Kevin Dale, Kalyan Sunkavalli, Micah K Johnson, Daniel Vlasic, Wojciech Matusik, and Hanspeter Pfister. Video face", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1095 }, { "text": "replacement. In SIGGRAPH, 2011. [177] Yuan Lin, Shengjin Wang, Qian Lin, and Feng Tang. Face swapping under large pose variations: A 3d model based approach. In ICME, 2012. [178] Jianke Zhu, Luc Van Gool, and Steven CH Hoi. Unsupervised face alignment by robust nonrigid mapping. In ICCV, 2009. [179] Saleh Mosaddegh, Loic Simon, and Fr\u00e9d\u00e9ric Jurie. Photorealistic face de-identification by aggregating donors\u2019 face", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1096 }, { "text": "components. In ACCV, 2015. [180] Ralph Gross, Iain Matthews, Jeffrey Cohn, Takeo Kanade, and Simon Baker. Multi-pie. IVC, 2010. [181] Stephen Milborrow, John Morkel, and Fred Nicolls. The muct landmarked face database. PRASA, 2010. [182] Yuval Nirkin, Iacopo Masi, Anh Tran Tuan, Tal Hassner, and Gerard Medioni. On face segmentation, face swapping,", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1097 }, { "text": "and face perception. In FG, 2018. [183] Xavier P Burgos-Artizzu, Pietro Perona, and Piotr Doll\u00e1r. Robust face landmark estimation under occlusion. In ICCV, 2013. [184] Jianmin Bao, Dong Chen, Fang Wen, Houqiang Li, and Gang Hua. Towards open-set identity preserving face synthesis. In CVPR, 2018. [185] Yandong Guo, Lei Zhang, Yuxiao Hu, Xiaodong He, and Jianfeng Gao. Ms-celeb-1m: A dataset and benchmark for", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1098 }, { "text": "large-scale face recognition. In ECCV, 2016. [186] Ning Zhang, Manohar Paluri, Yaniv Taigman, Rob Fergus, and Lubomir Bourdev. Beyond frontal faces: Improving person recognition using multiple cues. In CVPR, 2015. [187] Yuval Nirkin, Yosi Keller, and Tal Hassner. Fsgan: Subject agnostic face swapping and reenactment. In ICCV, 2019.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1099 }, { "text": "aware face swapping. CVPR, 2020. [190] Bingquan Zhu, Hao Fang, Yanan Sui, and Luming Li. Deepfakes for medical video de-identification: Privacy protection and diagnostic information preservation. In AAAI, 2020. [191] Renwang Chen, Xuanhong Chen, Bingbing Ni, and Yanhao Ge. Simswap: An efficient framework for high fidelity", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1100 }, { "text": "30 Pei and Zhang, et al. [193] Yuhan Wang, Xu Chen, Junwei Zhu, Wenqing Chu, Ying Tai, Chengjie Wang, Jilin Li, Yongjian Wu, Feiyue Huang, and Rongrong Ji. Hififace: 3d shape and semantic prior guided high fidelity face swapping. In IJCAI, 2021. [194] Yuval Nirkin, Yosi Keller, and Tal Hassner. Fsganv2: Improved subject agnostic face swapping and reenactment.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1101 }, { "text": "TPAMI, 2022. [195] Jiseob Kim, Jihoon Lee, and Byoung-Tak Zhang. Smooth-swap: a simple enhancement for face-swapping with smoothness. In CVPR, 2022. [196] Yixuan Li, Chao Ma, Yichao Yan, Wenhan Zhu, and Xiaokang Yang. 3d-aware face swapping. In CVPR, 2023. [197] Hao Zeng, Wei Zhang, Changjie Fan, Tangjie Lv, Suzhen Wang, Zhimeng Zhang, Bowen Ma, Lincheng Li, Yu Ding,", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1102 }, { "text": "and Xin Yu. Flowface: Semantic flow-guided shape-aware face swapping. In AAAI, 2023. [198] Fengyuan Liu, Lingyun Yu, Hongtao Xie, Chuanbin Liu, Zhiguo Ding, Quanwei Yang, and Yongdong Zhang. High fidelity face swapping via semantics disentanglement and structure enhancement. In ACM MM, 2023. [199] Yixuan Zhu, Wenliang Zhao, Yansong Tang, Yongming Rao, Jie Zhou, and Jiwen Lu. Stableswap: Stable face swapping", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1103 }, { "text": "in a shared and controllable latent space. TMM, 2024. [200] Wenliang Zhao, Yongming Rao, Weikang Shi, Zuyan Liu, Jie Zhou, and Jiwen Lu. Diffswap: High-fidelity and controllable face swapping via 3d-aware masked diffusion. In CVPR, 2023. [201] Kihong Kim, Yunho Kim, Seokju Cho, Junyoung Seo, Jisu Nam, Kychul Lee, Seungryong Kim, and KwangHee Lee.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1104 }, { "text": "Diffface: Diffusion-based face swapping with facial guidance. Pattern Recognition, 163:111451, 2025. [202] Sanoojan Baliah, Qinliang Lin, Shengcai Liao, Xiaodan Liang, and Muhammad Haris Khan. Realistic and efficient face swapping: A unified approach with diffusion models. In WACV, 2025. [203] Taewoo Kim, Geonsu Lee, Hyukgi Lee, Seongtae Kim, and Younggun Lee. Pixswap: High-resolution face swapping", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1105 }, { "text": "for effective reflection of identity via pixel-level supervision with synthetic paired dataset. In WACV, 2025. [204] Kaiwen Cui, Rongliang Wu, Fangneng Zhan, and Shijian Lu. Face transformer: Towards high fidelity and accurate face swapping. In CVPR, 2023. [205] Wei Cao, Tianyi Wang, Anming Dong, and Minglei Shu. Transfs: Face swapping using transformer. In FG, 2023.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1106 }, { "text": "flexible and extensible face-swapping framework. PR, 2023. [211] Chao Xu, Jiangning Zhang, Yue Han, Guanzhong Tian, Xianfang Zeng, Ying Tai, Yabiao Wang, Chengjie Wang, and Yong Liu. Designing one unified framework for high-fidelity face reenactment and swapping. In ECCV, 2022. [212] Diqiong Jiang, Dan Song, Ruofeng Tong, and Min Tang. Styleipsb: Identity-preserving semantic basis of stylegan for", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1107 }, { "text": "high fidelity face swapping. In CVPR, 2023. [213] Xiaohang Ren, Xingyu Chen, Pengfei Yao, Heung-Yeung Shum, and Baoyuan Wang. Reinforced disentanglement for face swapping without skip connection. In ICCV, 2023. [214] Yu Zhang, Hao Zeng, Bowen Ma, Wei Zhang, Zhimeng Zhang, Yu Ding, Tangjie Lv, and Changjie Fan. Flowface++:", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1108 }, { "text": "Explicit semantic flow-supervised end-to-end face swapping. arXiv, 2023. [215] Zhian Liu, Maomao Li, Yong Zhang, Cairong Wang, Qi Zhang, Jue Wang, and Yongwei Nie. Fine-grained face swapping via regional gan inversion. In CVPR, 2023. [216] Felix Rosberg, Eren Erdal Aksoy, Fernando Alonso-Fernandez, and Cristofer Englund. Facedancer: pose-and occlusion-", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1109 }, { "text": "aware high fidelity face swapping. In WACV, 2023. [217] Gege Gao, Huaibo Huang, Chaoyou Fu, Zhaoyang Li, and Ran He. Information bottleneck disentanglement for identity swapping. In CVPR, 2021. [218] Zhiliang Xu, Hang Zhou, Zhibin Hong, Ziwei Liu, Jiaming Liu, Zhizhi Guo, Junyu Han, Jingtuo Liu, Errui Ding, and", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1110 }, { "text": "Jingdong Wang. Styleswap: Style-based generator empowers robust face swapping. In ECCV, 2022. [219] Sahng-Min Yoo, Tae-Min Choi, Jae-Woo Choi, and Jong-Hwan Kim. Fastswap: A lightweight one-stage framework for real-time face swapping. In WACV, 2023. [220] Iryna Korshunova, Wenzhe Shi, Joni Dambre, and Lucas Theis. Fast face-swap using convolutional neural networks.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1111 }, { "text": "In ICCV, 2017. [221] Hyeongwoo Kim, Pablo Garrido, Ayush Tewari, Weipeng Xu, Justus Thies, Matthias Niessner, Patrick P\u00e9rez, Christian Richardt, Michael Zollh\u00f6fer, and Christian Theobalt. Deep video portraits. ACM TOG, 2018. [222] Hyeongwoo Kim, Mohamed Elgharib, Michael Zollh\u00f6fer, Hans-Peter Seidel, Thabo Beeler, Christian Richardt, and", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1112 }, { "text": "Deepfake Generation and Detection: A Benchmark and Survey 31 [223] Michail Christos Doukas, Stefanos Zafeiriou, and Viktoriia Sharmanska. Headgan: One-shot neural head synthesis and editing. In ICCV, 2021. [224] Kewei Yang, Kang Chen, Daoliang Guo, Song-Hai Zhang, Yuan-Chen Guo, and Weidong Zhang. Face2face \ud835\udf0c: Real-time", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1113 }, { "text": "high-resolution one-shot face reenactment. In ECCV, 2022. [225] Yue Gao, Yuan Zhou, Jinglu Wang, Xiao Li, Xiang Ming, and Yan Lu. High-fidelity and freely controllable talking head video generation. In CVPR, 2023. [226] Tina Behrouzi, Atefeh Shahroudnejad, and Payam Mousavi. Maskrenderer: 3d-infused multi-mask realistic face", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1114 }, { "text": "reenactment. PR, 2025. [227] Egor Zakharov, Aliaksandra Shysheya, Egor Burkov, and Victor Lempitsky. Few-shot adversarial learning of realistic neural talking head models. In CVPR, 2019. [228] Jiangning Zhang, Xianfang Zeng, Mengmeng Wang, Yusu Pan, Liang Liu, Yong Liu, Yu Ding, and Fan. Freenet: Multi-identity face reenactment. In CVPR, 2020.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1115 }, { "text": "strained gaze estimation in the wild. In ICCV, 2019. [231] Xucong Zhang, Yusuke Sugano, Mario Fritz, and Andreas Bulling. Appearance-based gaze estimation in the wild. In CVPR, 2015. [232] Bowen Zhang, Chenyang Qi, Pan Zhang, Bo Zhang, HsiangTao Wu, Dong Chen, Qifeng Chen, Yong Wang, and Fang Wen. Metaportrait: Identity-preserving talking head generation with fast personalized adaptation. In CVPR, 2023.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1116 }, { "text": "Liefeng Bo, and Xuelong Li. One-shot high-fidelity talking-head synthesis with deformable neural radiance field. In CVPR, 2023. [237] Stella Bounareli, Christos Tzelepis, Vasileios Argyriou, Ioannis Patras, and Georgios Tzimiropoulos. Stylemask: Disentangling the style space of stylegan2 for neural face reenactment. In FG, 2023.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1117 }, { "text": "face capture and reenactment of rgb videos. In CVPR, 2016. [245] Olivia Wiles, A Koepke, and Andrew Zisserman. X2face: A network for controlling face generation using images, audio, and pose codes. In ECCV, 2018. [246] Egor Zakharov, Aleksei Ivakhnenko, Aliaksandra Shysheya, and Victor Lempitsky. Fast bi-layer neural synthesis of", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1118 }, { "text": "one-shot realistic head avatars. In ECCV, 2020. [247] Sungjoo Ha, Martin Kersner, Beomsu Kim, Seokjun Seo, and Dongyoung Kim. Marionette: Few-shot face reenactment preserving identity of unseen targets. In AAAI, 2020. [248] Martin Cooke, Jon Barker, Stuart Cunningham, and Xu Shao. An audio-visual corpus for speech perception and", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1119 }, { "text": "automatic speech recognition. JASA, 2006. [249] Carolyn Richie, Sarah Warburton, and Megan Carter. Audiovisual database of spoken american english. LDC, 2009. [250] Hang Zhou, Yu Liu, Ziwei Liu, Ping Luo, and Xiaogang Wang. Talking face generation by adversarially disentangled audio-visual representation. In AAAI, 2019.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1120 }, { "text": "32 Pei and Zhang, et al. [251] Yang Zhou, Xintong Han, Eli Shechtman, Jose Echevarria, Evangelos Kalogerakis, and Dingzeyu Li. Makelttalk: speaker-aware talking-head animation. ACM TOG, 2020. [252] Xinya Ji, Hang Zhou, Kaisiyuan Wang, Wayne Wu, Chen Change Loy, Xun Cao, and Feng Xu. Audio-driven emotional", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1121 }, { "text": "video portraits. In CVPR, 2021. [253] Siddharth Gururani, Arun Mallya, Ting-Chun Wang, Rafael Valle, and Ming-Yu Liu. Space: Speech-driven portrait animation with controllable expression. In ICCV, 2023. [254] Wenxuan Zhang, Xiaodong Cun, Xuan Wang, Yong Zhang, Xi Shen, Yu Guo, Ying Shan, and Fei Wang. Sadtalker:", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1122 }, { "text": "Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation. In CVPR, 2023. [255] Ziqiao Peng, Haoyu Wu, Zhenbo Song, Hao Xu, Xiangyu Zhu, Jun He, Hongyan Liu, and Zhaoxin Fan. Emotalk: Speech-driven emotional disentanglement for 3d face animation. In ICCV, 2023.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1123 }, { "text": "and Jingdong Wang. Expressive talking head generation with granular audio-visual control. In CVPR, 2022. [262] Lingyun Yu, Hongtao Xie, and Yongdong Zhang. Multimodal learning for temporally coherent talking face generation with articulator synergy. TMM, 2021. [263] Chao Xu, Junwei Zhu, Jiangning Zhang, Yue Han, Wenqing Chu, Ying Tai, Chengjie Wang, Zhifeng Xie, and Yong Liu.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1124 }, { "text": "High-fidelity generalized emotional talking face generation with multi-modal emotion space learning. In CVPR, 2023. [264] Jiayu Wang, Kang Zhao, Shiwei Zhang, Yingya Zhang, Yujun Shen, Deli Zhao, and Jingren Zhou. Lipformer: High-fidelity and generalizable talking face generation with a pre-learned facial codebook. In CVPR, 2023.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1125 }, { "text": "speech-driven talking face generation with diffusion autoencoder. In ACM MM, 2023. [267] Zhentao Yu, Zixin Yin, Deyu Zhou, Duomin Wang, Finn Wong, and Baoyuan Wang. Talking head generation with probabilistic audio-to-visual diffusion priors. In ICCV, 2023. [268] Micha\u0142 Stypu\u0142kowski, Konstantinos Vougioukas, Sen He, Maciej Zi\u0119ba, Stavros Petridis, and Maja Pantic. Diffused", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1126 }, { "text": "heads: Diffusion models beat gans on talking-face generation. In WACV, 2024. [269] Bingyuan Zhang, Xulong Zhang, Ning Cheng, Jun Yu, Jing Xiao, and Jianzong Wang. Emotalker: Emotionally editable talking face generation via diffusion model. In ICASSP, 2024. [270] Sicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang, Chong Li, Zhenyu Zang, Yizhong Zhang, Xin Tong, and", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1127 }, { "text": "Baining Guo. Vasa-1: Lifelike audio-driven talking faces generated in real time. In NeurIPS, 2024. [271] Haotian Wang, Yuzhe Weng, Yueyan Li, Zilu Guo, Jun Du, Shutong Niu, Jiefeng Ma, Shan He, Xiaoyan Wu, Qiming Hu, et al. Emotivetalk: Expressive talking head generation through audio information decoupling and emotional", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1128 }, { "text": "video diffusion. In CVPR, 2025. [272] Yudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu, Hujun Bao, and Juyong Zhang. Ad-nerf: Audio driven neural radiance fields for talking head synthesis. In ICCV, 2021. [273] Shuai Shen, Wanhua Li, Zheng Zhu, Yueqi Duan, Jie Zhou, and Jiwen Lu. Learning dynamic facial radiance fields for", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1129 }, { "text": "few-shot talking head synthesis. In ECCV, 2022. [274] Dongze Li, Kang Zhao, Wei Wang, Bo Peng, Yingya Zhang, Jing Dong, and Tieniu Tan. Ae-nerf: Audio enhanced neural radiance field for few shot talking head synthesis. In AAAI, 2024. [275] Ziqiao Peng, Wentao Hu, Yue Shi, Xiangyu Zhu, Xiaomei Zhang, Hao Zhao, Jun He, Hongyan Liu, and Zhaoxin Fan.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1130 }, { "text": "Synctalk: The devil is in the synchronization for talking head synthesis. In CVPR\u201924, 2024. [276] Zhenhui Ye, Tianyun Zhong, Yi Ren, Jiaqi Yang, Weichuang Li, Jiawei Huang, Ziyue Jiang, Jinzheng He, Rongjie Huang, Jinglin Liu, et al. Real3d-portrait: One-shot realistic 3d talking portrait synthesis. In ICLR, 2024.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1131 }, { "text": "Deepfake Generation and Detection: A Benchmark and Survey 33 [279] Huayu Zhang, Yurui Ren, Yuanqi Chen, Ge Li, and Thomas H Li. Exploiting multiple guidance from 3dmm for face reenactment. In AAAIW, 2023. [280] Xiuzhe Wu, Pengfei Hu, Yang Wu, Xiaoyang Lyu, Yan-Pei Cao, Ying Shan, Wenming Yang, Zhongqian Sun, and", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1132 }, { "text": "Xiaojuan Qi. Speech2lip: High-fidelity speech to lip generation by learning from a short video. In ICCV, 2023. [281] Hui Fu, Zeqing Wang, Ke Gong, Keze Wang, Tianshui Chen, Haojie Li, Haifeng Zeng, and Wenxiong Kang. Mimic: Speaking style disentanglement for speech-driven 3d facial animation. In AAAI, 2024.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1133 }, { "text": "talking face generation with landmark and appearance priors. In CVPR, 2023. [287] Linrui Tian, Qi Wang, Bang Zhang, and Liefeng Bo. Emo: Emote portrait alive-generating expressive portrait videos with audio2video diffusion model under weak conditions. In ECCV, 2024. [288] Johanna Karras, Aleksander Holynski, Ting-Chun Wang, and Ira Kemelmacher-Shlizerman. Dreampose: Fashion", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1134 }, { "text": "image-to-video synthesis via stable diffusion. In ICCV, 2023. [289] Bernhard Kerbl, Georgios Kopanas, Thomas Leimk\u00fchler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM TOG, 2023. [290] Shuchang Zhou, Taihong Xiao, Yi Yang, Dieqiao Feng, Qinyao He, and Weiran He. Genegan: Learning object", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1135 }, { "text": "transfiguration and attribute subspace from unpaired data. In BMVC, 2017. [291] Youngjoo Jo and Jongyoul Park. Sc-fegan: Face editing generative adversarial network with user\u2019s sketch and color. In ICCV, 2019. [292] Zhenliang He, Wangmeng Zuo, Meina Kan, Shiguang Shan, and Xilin Chen. Attgan: Facial attribute editing by only", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1136 }, { "text": "changing what you want. TIP, 2019. [293] Yujun Shen, Jinjin Gu, Xiaoou Tang, and Bolei Zhou. Interpreting the latent space of gans for semantic face editing. In CVPR, 2020. [294] Xu Yao, Alasdair Newson, Yann Gousseau, and Pierre Hellier. A latent transformer for disentangled face editing in images and videos. In ICCV, 2021.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1137 }, { "text": "and editing. In CVPR, 2023. [301] Savas Ozkan, Mete Ozay, and Tom Robinson. Conceptual and hierarchical latent space decomposition for face editing. In ICCV, 2023. [302] Peng Zhou, Lingxi Xie, Bingbing Ni, and Qi Tian. Cips-3d++: End-to-end real-time high-resolution 3d-aware gans for gan inversion and stylization. TPAMI, 2023.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1138 }, { "text": "34 Pei and Zhang, et al. [307] Gyeongman Kim, Hajin Shim, Hyunsu Kim, Yunjey Choi, Junho Kim, and Eunho Yang. Diffusion video autoencoders: Toward temporally consistent face video editing via disentangled video encoding. In CVPR, 2023. [308] Weiqi Luo Wenmin Huang and Xiaochun Cao Jiwu Huang. Sdgan: Disentangling semantic manipulation for facial", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1139 }, { "text": "attribute editing. In AAAI, 2024. [309] Hao Zhang, Tianyuan DAI, Yanbo Xu, Yu-Wing Tai, and Chi-Keung Tang. Facednerf: Semantics-driven face recon- struction, prompt editing and relighting with diffusion models. In NeurIPS, 2024. [310] Kaiwen Jiang, Shu-Yu Chen, Feng-Lin Liu, Hongbo Fu, and Lin Gao. Towards high-quality and disentangled face", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1140 }, { "text": "editing in a 3d gan. TPAMI, 2025. [311] Yating Zeng, Xinpeng Zhang, and Guorui Feng. Secure reversible privacy protection for face multiple attribute editing. PR, 2025. [312] Bo Chen, Shoukang Hu, Qi Chen, Chenpeng Du, Ran Yi, Yanmin Qian, and Xie Chen. Gstalker: Real-time audio-driven talking face generation via deformable gaussian splatting. arXiv, 2024.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1141 }, { "text": "and manipulation. TMM, 2022. [316] Zhengzhe Liu, Xiaojuan Qi, and Philip HS Torr. Global texture enhancement for fake face detection in the wild. In CVPR, 2020. [317] Lingzhi Li, Jianmin Bao, Ting Zhang, Hao Yang, Dong Chen, Fang Wen, and Baining Guo. Face x-ray for more general face forgery detection. In CVPR, 2020.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1142 }, { "text": "synthetic video detector: From face or background manipulations to fully ai-generated content. In CVPR, 2025. [324] Yinglin Zheng, Jianmin Bao, Dong Chen, Ming Zeng, and Fang Wen. Exploring temporal coherence for more general video face forgery detection. In ICCV, 2021. [325] Alexandros Haliassos, Konstantinos Vougioukas, Stavros Petridis, and Maja Pantic. Lips don\u2019t lie: A generalisable", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1143 }, { "text": "and robust approach to face forgery detection. In CVPR, 2021. [326] Zhihao Gu, Yang Chen, Taiping Yao, Shouhong Ding, Jilin Li, and Lizhuang Ma. Delving into the local: Dynamic inconsistency learning for deepfake video detection. In AAAI, 2022. [327] Jongwook Choi, Taehoon Kim, Yonghyun Jeong, Seungryul Baek, and Jongwon Choi. Exploiting style latent flows for", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1144 }, { "text": "generalizing deepfake detection video detection. In CVPR, 2024. [328] Yuting Xu, Jian Liang, Lijun Sheng, and Xiao-Yu Zhang. Towards generalizable deepfake video detection with thumbnail layout and graph reasoning. IJCV, 2024. [329] Chunlei Peng, Zimin Miao, Decheng Liu, Nannan Wang, Ruimin Hu, and Xinbo Gao. Where deepfakes gaze at?", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1145 }, { "text": "spatial-temporal gaze inconsistency analysis for video face forgery detection. TIFS, 2024. [330] Changtao Miao, Zichang Tan, Qi Chu, Nenghai Yu, and Guodong Guo. Hierarchical frequency-assisted interactive networks for face manipulation detection. TIFS, 2022. [331] Zhiqing Guo, Zhenhong Jia, Liejun Wang, Dewang Wang, Gaobo Yang, and Nikola Kasabov. Constructing new", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1146 }, { "text": "backbone networks via space-frequency interactive convolution for deepfake detection. TIFS, 2023. [332] Chuangchuang Tan, Yao Zhao, Shikui Wei, Guanghua Gu, Ping Liu, and Yunchao Wei. Frequency-aware deepfake detection: Improving generalizability through frequency space domain learning. In AAAI, 2024.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1147 }, { "text": "Deepfake Generation and Detection: A Benchmark and Survey 35 [337] Juan Hu, Xin Liao, Jinwen Liang, Wenbo Zhou, and Zheng Qin. Finfer: Frame inference-based deepfake detection for high-visual-quality videos. In AAAI, 2022. [338] Xiao Guo, Xiaohong Liu, Zhiyuan Ren, Steven Grosz, Iacopo Masi, and Xiaoming Liu. Hierarchical fine-grained", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1148 }, { "text": "image forgery detection and localization. In CVPR, 2023. [339] Yuanhao Zhai, Tianyu Luan, David Doermann, and Junsong Yuan. Towards generic image manipulation detection with weakly-supervised self-consistency learning. In ICCV, 2023. [340] Jing Dong, Wei Wang, and Tieniu Tan. Casia image tampering detection evaluation database. In ChinaSIP, 2013.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1149 }, { "text": "that generalize. In ECCV, 2020. [346] Davide Cozzolino, Alessandro Pianese, Matthias Nie\u00dfner, and Luisa Verdoliva. Audio-visual person-of-interest deepfake detection. In CVPR, 2023. [347] Shruti Agarwal and Hany Farid. Detecting deep-fake videos from aural and oral dynamics. In CVPR, 2021. [348] Joel Frank, Thorsten Eisenhofer, Lea Sch\u00f6nherr, Asja Fischer, Dorothea Kolossa, and Thorsten Holz. Leveraging", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1150 }, { "text": "frequency analysis for deep fake image recognition. In ICML, 2020. [349] Iacopo Masi, Aditya Killekar, Royston Marian Mascarenhas, Shenoy Pratik Gurudatt, and Wael AbdAlmageed. Two-branch recurrent network for isolating deepfakes in videos. In ECCV, 2020. [350] Ying Guo and Cheng Zhen. Controllable guide-space for generalizable face forgery detection. In ICCV, 2023.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1151 }, { "text": "layers. In CVPR, 2015. [360] Hung-Jen Chen, Ka-Ming Hui, Szu-Yu Wang, Li-Wu Tsao, Hong-Han Shuai, and Wen-Huang Cheng. Beautyglow: On-demand makeup transfer framework with reversible generative network. In CVPR, 2019. [361] Yu Deng, Jiaolong Yang, Sicheng Xu, Dong Chen, Yunde Jia, and Xin Tong. Accurate 3d face reconstruction with", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1152 }, { "text": "weakly-supervised learning: From single image to image set. In CVPR, 2019. [362] Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou. Arcface: Additive angular margin loss for deep face recognition. In CVPR, 2019. [363] Madhav Agarwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar. Audio-visual face reenactment.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1153 }, { "text": "In WACV, 2023. [364] Peng Tang, Huihuang Zhao, Weiliang Meng, and Yaonan Wang. One-shot motion talking head generation with audio-driven model. ESWA, 2025. [365] Seunggyu Chang, Gihoon Kim, and Hayeon Kim. Hairnerf: Geometry-aware image synthesis for hairstyle transfer. In ICCV, 2023. [366] Qiyao Deng, Jie Cao, Yunfan Liu, Qi Li, and Zhenan Sun. r-face: Reference guided face component editing. PR, 2024.", "source": "Deepfake Generation and Detection Survey", "year": 2024, "url": "https://arxiv.org/abs/2403.17881", "id": 1154 }, { "text": "Frequency-Aware Deepfake Detection: Improving Generalizability through Frequency Space Learning Chuangchuang Tan1,2, Yao Zhao1,2*, Shikui Wei1,2, Guanghua Gu3,4, Ping Liu5, Yunchao Wei1,2 1Institute of Information Science, Beijing Jiaotong University 2Beijing Key Laboratory of Advanced Information Science and Network Technology", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1155 }, { "text": "3School of Information Science and Engineering, Yanshan University 4Hebei Key Laboratory of Information Transmission and Signal Processing 5Center for Frontier AI Research, IHPC, A*STAR, Singapore {tanchuangchuang, yzhao, shkwei}@bjtu.edu.cn, guguanghua@ysu.edu.cn, pino.pingliu@gmail.com, wychao1987@gmail.com", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1156 }, { "text": "Abstract This research addresses the challenge of developing a uni- versal deepfake detector that can effectively identify un- seen deepfake images despite limited training data. Exist- ing frequency-based paradigms have relied on frequency- level artifacts introduced during the up-sampling in GAN pipelines to detect forgeries. However, the rapid advance-", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1157 }, { "text": "ments in synthesis technology have led to specific artifacts for each generation model. Consequently, these detectors have exhibited a lack of proficiency in learning the fre- quency domain and tend to overfit to the artifacts present in the training data, leading to suboptimal performance on unseen sources. To address this issue, we introduce a novel", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1158 }, { "text": "frequency-aware approach called FreqNet, centered around frequency domain learning, specifically designed to enhance the generalizability of deepfake detectors. Our method forces the detector to continuously focus on high-frequency infor- mation, exploiting high-frequency representation of features", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1159 }, { "text": "across spatial and channel dimensions. Additionally, we in- corporate a straightforward frequency domain learning mod- ule to learn source-agnostic features. It involves convolu- tional layers applied to both the phase spectrum and am- plitude spectrum between the Fast Fourier Transform (FFT) and Inverse Fast Fourier Transform (iFFT). Extensive experi-", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1160 }, { "text": "mentation involving 17 GANs demonstrates the effectiveness of our proposed method, showcasing state-of-the-art perfor- mance (+9.8%) while requiring fewer parameters. The code is available at https://github.com/chuangchuangtan/FreqNet- DeepfakeDetection. Introduction The proliferation of Generative Adversarial Networks", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1161 }, { "text": "(GANs) (Goodfellow et al. 2014; Karras et al. 2018, 2019) has significantly simplified the generation of lifelike syn- thetic images, resulting in an alarming surge in the preva- lence of forgeries that are virtually indistinguishable from authentic images to the human visual system. This escalating", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1162 }, { "text": "*Corresponding author Copyright \u00a9 2024, Association for the Advancement of Artificial Intelligence (www.aaai.org). All rights reserved. Conv High-Freq Representation Conv Residual block Conv Residual block real or fake? image freq-level artifact CNN classifier real or fake? image high-freq component Classifier with frequency learning", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1163 }, { "text": "Conv FFT Conv Conv iFFT Conv (a) Conventional frequency-based methods (b) FreqNet: Frequency space learning Network phase spectrum amplitude spectrum Figure 1: Frequency space learning network. (a) The traditional studies are usually limited to developing frequency-level artifacts. (b) Distinguishing itself from prior", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1164 }, { "text": "frequency-based research, our approach shifts its focus to the frequency-related attributes of the features within the detector. This novel perspective includes a continuous em- phasis on high-frequency details within the classifier, which capitalizes on the enriched depiction of high-frequency fea-", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1165 }, { "text": "ture map components spanning both spatial and channel di- mensions. Additionally, our strategy introduces a trainable layer embedded within the frequency domain, facilitating the acquisition of source-agnostic features. trend poses potential, unpredictable societal repercussions. In response, a multitude of deepfake detection mechanisms", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1166 }, { "text": "have been conceived (Frank et al. 2020; Li et al. 2021), with specific emphasis on detecting facial forgeries. How- ever, the majority of existing forgery detection techniques suffer from a fundamental limitation: they are constrained to the same domain during both their training and evalua- tion phases. This limitation severely hampers their ability to", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1167 }, { "text": "to create a universal detector capable of effectively identify- ing deepfake images even when faced with limited training data, a necessity given the continual emergence of increas- ingly sophisticated synthesis technologies. Recently, prior investigations (Frank et al. 2020; Durall et al. 2020) have", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1168 }, { "text": "substantiated the efficacy of frequency artifacts in the realm of deepfake detection, notably in the context of facial de- tection (Qian et al. 2020; Luo et al. 2021). The findings of (Frank et al. 2020; Durall et al. 2020) have unveiled the presence of significant artifacts within the frequency domain", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1169 }, { "text": "stemming from the upsampling operations within GAN ar- chitectures. Consequently, this revelation has spurred the de- velopment of numerous frequency-based methods aimed at detecting images with pronounced frequency-related char- acteristics. Nonetheless, owing to the extraordinary advancements in synthesis technology, an increasing array of distinctive", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1170 }, { "text": "frequency-level artifact representations have emerged. In Figure 2, we present the mean Fast Fourier Transform (FFT) (Cooley et al. 1969) spectrum of images sampled from var- ious sources. This mean spectrum computation involves av- eraging over 2,000 images, following the methodology de- tailed in (Frank et al. 2020). Notably, the results exhibit dis-", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1171 }, { "text": "cernible differences in artifact characteristics across differ- ent GANs, further accentuated by variations within the same GAN architecture when training on dissimilar datasets. The frequency attributes of images indeed possess the capacity to unveil distinctions between real and generated images.", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1172 }, { "text": "However, they exhibit limitations in terms of generalization across a diverse range of sources. The CNN classifier, when trained on a specific source (e.g., StyleGAN), tends to overfit to the particular pattern present within the training data. As a result, this classifier often falters when faced with unseen", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1173 }, { "text": "synthesis models such as CycleGAN and BigGAN. To surmount this challenge, the primary approach entails the development of a robust classifier specifically designed for frequency representations. (Jeong et al. 2022c), in ad- dressing this issue, chooses to disregard frequency-level ar- tifacts in images by devising a frequency-level perturbation", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1174 }, { "text": "generator. However, this solution introduces complexity and incurs considerable computational expenses. In the realm of deepfake detection, it is imperative to factor in the notion of the detector learning within the frequency domain. We deliberately refrain from directly utilizing the frequency in-", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1175 }, { "text": "formation as the artifact representation to train a CNN clas- sifier. Instead, we strategically compel the detector to ac- quire its understanding within the frequency space. This nu- anced strategy holds the key to achieving a more generaliz- able deepfake detection framework. Drawing upon intuitive insights, we introduce a novel and", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1176 }, { "text": "lightweight approach named FreqNet, which integrates fre- quency domain learning into a lightweight CNN classifier, aimed at enhancing the generalization capabilities of the de- tector. Diverging from existing frequency-based studies that predominantly focus on the frequency domain of images, the principal innovation of our FreqNet method resides in", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1177 }, { "text": "the simultaneous development of both frequency-domain in- formation derived from images and the features extracted by CNN model. This distinctive dual approach empowers Apple Dataset CycleGAN-Apple Lsun-Bedroom Dataset StyleGAN-Bedroom Lsun-Church Dataset StyleGAN2-Church BigGAN ImageNet Dataset Zebra Dataset", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1178 }, { "text": "CycleGAN-Zebra Face Deepfake Figure 2: Frequency analysis on various sources. This mean FFT spectrum computation involves averaging over 2,000 images, following the methodology detailed in (Frank et al. 2020). the detector to acquire proficiency in the frequency domain, thereby diminishing its reliance on the specific frequency", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1179 }, { "text": "patterns present in the training source. Specifically, our approach introduces two critical mod- ules: the high-frequency representation and the frequency convolutional layer, each meticulously designed to facilitate frequency space learning. The first module serves to com- pel the detector to consistently prioritize high-frequency in-", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1180 }, { "text": "formation, augmenting its sensitivity to significant details. Additionally, to capture broader forgery indicators within the frequency domain, we incorporate a frequency convo- lutional layer, which effectively diminishes the reliance on source-specific characteristics. By virtue of this frequency domain-based learning strategy, our proposed FreqNet re-", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1181 }, { "text": "markably extends its generalizability to previously unseen sources with few parameters. To comprehensively assess the extent of its generalization capabilities, we conduct extensive simulations using an ex- tensive image database generated by 17 distinct models 1. Despite its modest scale, FreqNet, boasting 1.9 million pa-", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1182 }, { "text": "rameters, significantly outperforms the current state-of-the- art model boasting 304 million parameters, demonstrating a remarkable improvement of 9.8%. Our paper makes the following contributions: \u2022 We present a novel frequency space learning network, FreqNet, to achieve generalizable deepfake detection.", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1183 }, { "text": "Our approach strategically incorporates frequency do- main learning within a CNN classifier, resulting in a sig- nificant enhancement of the detector\u2019s ability to general- ize across diverse scenarios. \u2022 We utilize convolutional layers on both the phase spec- trum and the amplitude spectrum as a deliberate strat-", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1184 }, { "text": "egy to capture broader forgery indicators within the fre- quency domain. This allows us to enhance the detector\u2019s capability to identify a broader range of artifacts. \u2022 The proposed lightweight FreqNet, consisting of a mere 1.9 million parameters, impressively outperforms the current state-of-the-art model featuring 304 million pa-", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1185 }, { "text": "rameters, attributed mainly to its utilization of frequency domain learning. Related Work In this section, we present a concise survey of deepfake de- tection methodologies, categorizing them into two primary classes: image-based detection and frequency-based detec- tion. Image-based Deepfake Detection", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1186 }, { "text": "There has been a significant effort in the field of forgery detection, with many studies focusing on leveraging spa- tial information from images. Rossler et al.(Rossler et al. 2019) utilize images to train the Xception (Chollet 2017) to achieve fake face image detection. Other image-based de- tection methods have developed specific artifact detection in", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1187 }, { "text": "distinct facial regions, such as eyes(Li et al. 2018) and lips (Haliassos et al. 2021). Chai et al.(Chai et al. 2020) adopt limited receptive fields to identify patches that render im- ages detectable, highlighting the importance of specific local features. With the emergence of deepfake technology, efforts", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1188 }, { "text": "are being made to enhance the generalization ability of de- tectors, particularly for unseen data. Yu et al.(Yu et al. 2020) introduce artifacts from the camera imaging process. Fur- thermore, various methods (Wang et al. 2020, 2021; Chen et al. 2022; Cao et al. 2022; He et al. 2021; Shiohara et al.", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1189 }, { "text": "2022) aim to enrich the diversity of training data, employing techniques such as data augmentation, adversarial training, reconstruction, and blending images. CDDB (Li et al. 2023) adopts incremental learning to achieve continual deepfake detection. Notably, recent works by Ojha et al.(Ojha et al. 2023) and Tan et al.(Tan et al. 2023) employ the feature map", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1190 }, { "text": "and gradients as the general representation, respectively. Frequency-based Deepfake Detection The research by (Frank et al. 2020; Durall et al. 2020), which highlights the effectiveness of frequency artifacts in the domain of deepfake detection, has significantly influ- enced subsequent studies. These findings have led many im-", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1191 }, { "text": "age forgery detectors to shift their attention toward capturing unique patterns within the frequency domain. The work by Masi et al. (Masi et al. 2020) stands out as it meticulously investigates artifacts present in both the color space and the frequency domains, while F 3-Net (Qian et al. 2020) sug-", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1192 }, { "text": "gests using the discrepancy of frequency statistics between real and forged images as a means to differentiate face im- age manipulations. An adaptive frequency features learning is designed by FDFL (Li et al. 2021) to mine subtle arti- facts from the frequency domain, enabling the detection of forged images. In parallel, Luo et al.(Luo et al. 2021) im-", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1193 }, { "text": "prove the generalization performance through the integra- tion of multiple high-frequency features, thereby fortifying the robustness of the forgery detection process. Moreover, ADD (Woo et al. 2022) incorporates two meticulously de- signed distillation modules to emphasize the significance of frequency information, incorporating frequency attention", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1194 }, { "text": "distillation and multi-view attention distillation. Recently, frequency-based studies have been proposed for generalized detection. BiHPF (Jeong et al. 2022a) emphasizes amplify- ing artifact magnitudes through the utilization of dual high- pass filters. The FreGAN model, introduced by (Jeong et al.", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1195 }, { "text": "2022c) ingeniously mitigates the impact of frequency-level artifacts through the deployment of frequency-level pertur- bation maps. (Wang et al. 2023) introduces dynamic graph learning to exploit the relation-aware features in spatial and frequency domains. Methodology In this section, we present our FreqNet technique, a univer-", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1196 }, { "text": "sal deepfake detection method for generalizable deepfake detection. We illustrate the overall architecture of FreqNet in Figure 3, leveraging frequency domain learning to miti- gate source-specific dependencies. Problem Definition Our primary focus lies within the realm of Generalizable Deepfake Detection. Our objective is to construct a universal", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1197 }, { "text": "detector capable of accurately identifying deepfake images even when confronted with constrained training sources. In this context, let us consider a real-world image scenario de- noted as X sampled from n different sources: X = {X1, X2, ..., Xi, ..., Xn}, Xi = {xi j, yj}Ni j=1, (1) where Ni represents the number of images originating from", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1198 }, { "text": "the ith source Xi, xi j is the jth image of Xi. Each image is labeled with y, indicating whether it belongs to the category of \u201dreal\u201d(y = 0) or \u201dfake\u201d (y = 1). Here we train a binary classifier D(\u00b7), utilizing the training source Xi: Di = arg min \u03b8 l(D(Xi; \u03b8), y), (2) where l() denote the loss function. Our overarching goal is to", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1199 }, { "text": "design a detector that is trained on the data originating from Xi, but demonstrates strong performance when confronted with images coming from previously unseen sources, de- noted as Xt. This generalizability across unseen sources is a crucial objective of our detector. Overall Architecture With the primary objective of bolstering generalizability to", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1200 }, { "text": "unseen sources, we have devised a frequency domain learn- ing network to effectively enhance deepfake detection capa- bilities. The comprehensive architecture of the FreqNet ap- proach is depicted in Figure 3. Within our methodology, we introduce practical and compact frequency learning plugin modules designed to compel the CNN classifier to operate", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1201 }, { "text": "High-Frequency Representation of Image(HFRI) Residual Conv block FFT iFFT Frequency Conv Layer(FCL) FFT iFFT Phase Spectrum Conv Amplitude Spectrum Conv spectrum learning Residual Conv block FFT C dimension W dimension H dimension FFT Spatial (W,H) dim iFFT iFFT High-Frequency Representation of Feature across Spatial (HFRFS)", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1202 }, { "text": "Channel (C) dim High-Frequency Representation of Feature across Channel (HFRFC) HFRI Image HFRFC FCL HFRFS HFRFC FCL Residual Conv block HFRFS FCL HFRFS Conv Conv FCL Residual Conv block Global avg pooling FC layer FreqNet: Frequency space learning Network Overall Network (b) HFRF Block (a) HFRI Block", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1203 }, { "text": "(c) FCL Block Figure 3: Architecture of FreqNet for generalizable deepfake detection. To augment the capacity for generalization, our FreqNet focuses on the enhancement of frequency spectrum information, prioritizing frequency domain learning within the classifier, consisting of (a) High-Frequency Representation of Image(HFRI), (b) High-Frequency Representation of Feature(HFRF), and", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1204 }, { "text": "(c) Frequency Conv Layer(FCL). seamless integration into the CNN classifier. This architec- tural innovation has culminated in the creation of a novel and lightweight detector named FreqNet, incorporating the fre- quency learning plugins alongside a limited number of CNN layers. High-Frequency Representation", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1205 }, { "text": "of Images High- frequency artifacts have been recognized as valuable indicators for distinguishing between real and fake images (Qian et al. 2020; Luo et al. 2021). Consequently, we lever- age this valuable insight by adopting the high-frequency components of images as the input for the detector in our", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1206 }, { "text": "approach. In the case of each training image denoted as x \u2208RW \u00d7H\u00d73, our initial step involves converting it into the frequency domain using the Fast Fourier Transform (FFT). Subsequently, we proceed to extract the high-frequency components, represented as fh, through the application of a high-pass filter denoted as Bh:", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1207 }, { "text": "fh = Bh(F(x)), (3) where F denotes FFT, fh \u2208RW \u00d7H\u00d73 denotes the fre- quency repersentation of images. The zero-frequency is shifted to the center. The high-pass filter Bh(\u00b7) can be de- fined by: Bh(fi,j) = \u001afi,j, otherwise, 0, if |i| < W/4, |j| < H/4 (4) where the center of the image is adopted as the origin.", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1208 }, { "text": "Following the extraction of high-frequency components, we proceed to transform this frequency information back to image space: xh = IF(fh), (5) where the IF denotes inverse Fast Fourier Transform (iFFT). The result of this transformation, denoted as xh, represents the high-frequency components within the image", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1209 }, { "text": "space. This process ensures that our focus remains on the pertinent high-frequency information contained within the image. High-Frequency Representation of Feature Indeed, the varying GAN architectures incorporate distinct frequency patterns, which can potentially lead the detector to overfit to the specifics of the training source. To mitigate this issue and", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1210 }, { "text": "enhance the detector\u2019s generalization capacity, we adopt a strategy that involves compelling the detector to consistently prioritize and focus on high-frequency information within the feature space. By emphasizing the importance of high- frequency cues in the feature space, we effectively counter- act the overfitting tendencies, promoting a more robust and", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1211 }, { "text": "adaptable deepfake detection capability. Specifically, for output of the kth convolutional layer de- noted as M k \u2208RH\u00d7W \u00d7C, we transform it to frequency space across spatial (W, H) and the channel dimension C using the Fast Fourier Transform (FFT), respectively. The zero frequency component is moved to the center. Subse-", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1212 }, { "text": "quently, we apply a high-pass filter Bh to extract the high- frequency information from this transformed representation. Finally, we reverse this frequency transformation, convert- ing the extracted high-frequency information back to the fea- ture space: M k h(dim) = ( IFW,H(Bh(FW,H(M k))),ifdim = (W, H)", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1213 }, { "text": "IFC(Bh(FC(M k))),ifdim = C (6) where M k h(W, H), M k h(C) represent the high-frequency components of feature maps M k acorss spatial (W, H) and channel C dimension in feature space, respectively. We im- plement two distinct high-frequency component extractors, each targeting different CNN layers within our method. This", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1214 }, { "text": "ing the overall sensitivity to high-frequency cues. Frequency Convolutional Layer Many existing frequency-based approaches follow a paradigm of ex- tracting frequency information from images and employing it to train a CNN classifier. However, this approach often results in the detector overfitting to the specifics of the", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1215 }, { "text": "training source, leading to suboptimal performance when faced with previously unseen sources. In contrast, our approach not only employs the frequency information as the artifact representation. We also introduce frequency space learning as a strategy to significantly enhance the generalization ability of the detector.", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1216 }, { "text": "Specifically, within our approach, the feature maps from the convolutional layers are initially transformed from the feature space to the frequency domain. Following this trans- formation, we apply convolutional layers on both the phase spectrum and the amplitude spectrum, thereby enabling the detector to learn within the frequency space. Subsequently,", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1217 }, { "text": "the learned spectrum information is transformed back to the feature space using the inverse Fast Fourier Transform. This comprehensive process effectively emphasizes the represen- tation ability of the detector within the frequency domain, enhancing its sensitivity to critical features present in this", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1218 }, { "text": "space. Let\u2019s consider the given feature maps M k \u2208RW \u00d7H\u00d7C from kth CNN layer. The learning process within the fre- quency space can be formally defined as follows: f = fam + fphi = FW,H(M k) g fam = Lconv(fam) f fph = Lconv(fph) g M k = IFW,H(g fam + f fphi) (7) where FW,H, IFW,H denote FFT and iFFT, fam, fph are", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1219 }, { "text": "amplitude spectrum and phase spectrum of feature maps M k, and Lconv denotes a CNN layer, g fam, f fph are the learned amplitude spectrum and phase spectrum of feature maps M k. Subsequently, The feature map g M k learned in spectrum space can be calculated by iFFT. As a result, the comprehensive training procedure for Fre-", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1220 }, { "text": "qNet can be formally defined as: Dfreq = arg min \u03b8 l(Dfreq(xh; \u03b8), y), (8) where l() denotes the standard cross entropy loss, and Dfreq is our frequency sapce learning network. Our FreqNet har- nesses spectrum learning to accomplish domain-invariant deepfake detection. Within our approach, we have meticu-", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1221 }, { "text": "lously designed two key modules: the high-frequency repre- sentation module and the frequency convolutional layer. Im- portantly, this work places emphasis on training our detec- tor using a constrained amount of training data, followed by comprehensive evaluations in the challenging wild scenes, encompassing a diverse set of 17 GAN models.", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1222 }, { "text": "Experiments In this section, we provide a comprehensive evaluation of the FreqNet. We cover various aspects, including datasets, im- plementation details, deepfake detection performance, and more details to be described. Further elaborations on each of these aspects will be presented to offer a comprehensive", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1223 }, { "text": "understanding of the capabilities and effectiveness of Fre- qNet. Datasets Training set. To ensure a consistent basis for compari- son, we employ the training set of ForenSynths (Wang et al. 2020) to train the detectors, aligning with baselines (Wang et al. 2020; Jeong et al. 2022a,c). The training set consists", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1224 }, { "text": "of 20 distinct categories, each comprising 18,000 synthetic images generated using ProGAN, alongside an equal num- ber of real images sourced from the LSUN dataset. In line with previous research (Jeong et al. 2022a,c), we adopt spe- cific 1-class, 2-class, and 4-class training settings, denoted as", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1225 }, { "text": "(horse), (chair, horse), (car, cat, chair, horse), respectively. Real-world Scene Test set. To assess the generalization ability of the proposed method on the real-world scene, we adopt various images and diverse GAN models. Firstly, we employ the test set of ForenSynths for evaluation. It includes", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1226 }, { "text": "fake images generated by 8 generation model 2. The real im- ages are sampled from 6 datasets 3. Additionally, to replicate the unpredictability of wild scenes, we extend our evalua- tion by collecting images generated by 9 additional GANs 4. There are 36K test images, with equal numbers of real and", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1227 }, { "text": "fake images. Simultaneously, we curate a dedicated face test set com- prising 20,000 real images sourced from Celeba-HQ (Kar- ras et al. 2018), and 60,000 fake face images from ProGAN (Karras et al. 2018), StyleGAN (Karras et al. 2019), and StyleGAN2 (Karras et al. 2020). Implementation Details We design a lightweight CNN classifier, employing residual", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1228 }, { "text": "convolutional blocks without pretraining. During the train- ing process, we utilize the Adam optimizer (Kingma et al. 2015) with an initial learning rate of 2 \u00d7 10\u22122. The batch size is set at 32, and we train the model for 100 epochs. A learning rate decay strategy is employed, reducing the learn- ing rate by twenty percent after every ten epochs. Consistent", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1229 }, { "text": "with established baselines (Jeong et al. 2022a,c), we utilize the average precision score (A.P.) and accuracy (Acc.) as the primary evaluation metrics to gauge the effectiveness of our proposed method. These metrics provide a comprehen- sive assessment of the performance of our approach against 2ProGAN (Karras et al. 2018), StyleGAN (Karras et al. 2019),", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1230 }, { "text": "StyleGAN2 (Karras et al. 2020), BigGAN (Brock et al. 2018), Cy- cleGAN (Zhu et al. 2017), StarGAN (Choi et al. 2018), GauGAN (Park et al. 2019) and Deepfake (Rossler et al. 2019) 3LSUN (Yu et al. 2015), ImageNet (Russakovsky et al. 2015), CelebA (Liu et al. 2015), CelebA-HQ (Karras et al. 2018), COCO", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1231 }, { "text": "(Lin et al. 2014), and FaceForensics++ (Rossler et al. 2019) 4AttGAN(He et al. 2019), BEGAN(Berthelot et al. 2017), CramerGAN(Bellemare et al. 2017), InfoMaxGAN(Lee et al. 2021), MMDGAN(Li et al. 2017), RelGAN(Nie et al. 2019), S3GAN(Lu\u02c7ci\u00b4c et al. 2019), SNGAN(Miyato et al. 2018), and STGAN(Liu et al. 2019)", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1232 }, { "text": "Methods Settings Test Models Input #n ProGAN StyleGAN StyleGAN2 BigGAN CycleGAN StarGAN GauGAN Deepfake Mean Acc. A.P. Acc. A.P. Acc. A.P. Acc. A.P. Acc. A.P. Acc. A.P. Acc. A.P. Acc. A.P. Acc. A.P. Wang(2020) Img 1 50.4 63.8 50.4 79.3 68.2 94.7 50.2 61.3 50.0 52.9 50.0 48.2 50.3 67.6 50.1 51.5 52.5", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1233 }, { "text": "64.9 Frank(2020) Freq 1 78.9 77.9 69.4 64.8 67.4 64.0 62.3 58.6 67.4 65.4 60.5 59.5 67.5 69.1 52.4 47.3 65.7 63.3 Durall(2020) Freq 1 85.1 79.5 59.2 55.2 70.4 63.8 57.0 53.9 66.7 61.4 99.8 99.6 58.7 54.8 53.0 51.9 68.7 65.0 F3Net(2020) Freq 1 96.9 99.9 86.3 99.8 80.5 99.8 66.6 72.2 76.7 84.0 99.1 100.0", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1234 }, { "text": "59.1 60.6 61.2 82.3 78.3 87.3 BiHPF(2022a) Freq 1 82.5 81.4 68.0 62.8 68.8 63.6 67.0 62.5 75.5 74.2 90.1 90.1 73.6 92.1 51.6 49.9 72.1 72.1 FrePGAN(2022c) Img 1 95.5 99.4 80.6 90.6 77.4 93.0 63.5 60.5 59.4 59.9 99.6 100.0 53.0 49.1 70.4 81.5 74.9 79.3 LGrad (2023) Grad 1 99.4 99.9 96.0 99.6 93.8 99.4", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1235 }, { "text": "79.5 88.9 84.7 94.4 99.5 100.0 70.9 81.8 66.7 77.9 86.3 92.7 Ojha (2023) Fea 1 99.1 100.0 77.2 95.9 69.8 95.8 94.5 99.0 97.1 99.9 98.0 100.0 95.7 100.0 82.4 91.7 89.2 97.8 FreqNet Freq 1 98.0 99.9 92.0 98.7 89.5 97.9 85.5 93.1 96.1 99.1 94.2 98.4 91.8 99.6 69.8 94.4 89.6 97.6 Wang(2020) Img 2 64.6 92.7", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1236 }, { "text": "52.8 82.8 75.7 96.6 51.6 70.5 58.6 81.5 51.2 74.3 53.6 86.6 50.6 51.5 57.3 79.6 Frank(2020) Freq 2 85.7 81.3 73.1 68.5 75.0 70.9 76.9 70.8 86.5 80.8 85.0 77.0 67.3 65.3 50.1 55.3 75.0 71.2 Durall(2020) Freq 2 79.0 73.9 63.6 58.8 67.3 62.1 69.5 62.9 65.4 60.8 99.4 99.4 67.0 63.0 50.5 50.2 70.2 66.4 F3Net(2020)", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1237 }, { "text": "Freq 2 97.9 100.0 84.5 99.5 82.2 99.8 65.5 73.4 81.2 89.7 100.0 100.0 57.0 59.2 59.9 83.0 78.5 88.1 BiHPF(2022a) Freq 2 87.4 87.4 71.6 74.1 77.0 81.1 82.6 80.6 86.0 86.6 93.8 80.8 75.3 88.2 53.7 54.0 78.4 79.1 FrePGAN(2022c) Img 2 99.0 99.9 80.8 92.0 72.2 94.0 66.0 61.8 69.1 70.3 98.5 100.0 53.1 51.0", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1238 }, { "text": "62.2 80.6 75.1 81.2 LGrad (2023) Grad 2 99.8 100.0 94.8 99.7 92.4 99.6 82.5 92.4 85.9 94.7 99.7 99.9 73.7 83.2 60.6 67.8 86.2 92.2 Ojha (2023) Fea 2 99.7 100.0 78.8 97.4 75.4 96.7 91.2 99.0 91.9 99.8 96.3 99.9 91.9 100.0 80.0 89.4 88.1 97.8 FreqNet Freq 2 99.6 100.0 90.4 98.9 85.8 98.1 89.0 96.0 96.7", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1239 }, { "text": "99.8 97.5 100.0 88.0 98.8 80.7 92.0 91.0 97.9 Wang(2020) Img 4 91.4 99.4 63.8 91.4 76.4 97.5 52.9 73.3 72.7 88.6 63.8 90.8 63.9 92.2 51.7 62.3 67.1 86.9 High-Freq Freq 4 98.9 100.0 74.4 98.3 68.8 97.3 75.2 92.1 71.0 87.9 92.7 100.0 75.5 86.5 57.0 74.9 76.7 92.1 Frank(2020) Freq 4 90.3 85.2 74.5 72.0", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1240 }, { "text": "73.1 71.4 88.7 86.0 75.5 71.2 99.5 99.5 69.2 77.4 60.7 49.1 78.9 76.5 Durall(2020) Freq 4 81.1 74.4 54.4 52.6 66.8 62.0 60.1 56.3 69.0 64.0 98.1 98.1 61.9 57.4 50.2 50.0 67.7 64.4 F3Net(2020) Freq 4 99.4 100.0 92.6 99.7 88.0 99.8 65.3 69.9 76.4 84.3 100.0 100.0 58.1 56.7 63.5 78.8 80.4 86.2 BiHPF(2022a)", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1241 }, { "text": "Freq 4 90.7 86.2 76.9 75.1 76.2 74.7 84.9 81.7 81.9 78.9 94.4 94.4 69.5 78.1 54.4 54.6 78.6 77.9 FrePGAN(2022c) Img 4 99.0 99.9 80.7 89.6 84.1 98.6 69.2 71.1 71.1 74.4 99.9 100.0 60.3 71.7 70.9 91.9 79.4 87.2 LGrad (2023) Grad 4 99.9 100.0 94.8 99.9 96.0 99.9 82.9 90.7 85.3 94.0 99.6 100.0 72.4 79.3", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1242 }, { "text": "58.0 67.9 86.1 91.5 Ojha (2023) Fea 4 99.7 100.0 89.0 98.7 83.9 98.4 90.5 99.1 87.9 99.8 91.4 100.0 89.9 100.0 80.2 90.2 89.1 98.3 FreqNet Freq 4 99.6 100.0 90.2 99.7 88.0 99.5 90.5 96.0 95.8 99.6 85.7 99.8 93.4 98.6 88.9 94.4 91.5 98.5 Table 1: Cross-model performance on the test set of ForenSynths(Wang et al. 2020). Bold and underline represent the best and", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1243 }, { "text": "second-best performance, respectively. Method AttGAN BEGAN CramerGAN InfoMaxGAN MMDGAN RelGAN S3GAN SNGAN STGAN Mean Acc. A.P. Acc. A.P. Acc. A.P. Acc. A.P. Acc. A.P. Acc. A.P. Acc. A.P. Acc. A.P. Acc. A.P. Acc. A.P. Wang(2020) 51.1 83.7 50.2 44.9 81.5 97.5 71.1 94.7 72.9 94.4 53.3 82.1 55.2 66.1 62.7", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1244 }, { "text": "90.4 63.0 92.7 62.3 82.9 F3Net(2020) 85.2 94.8 87.1 97.5 89.5 99.8 67.1 83.1 73.7 99.6 98.8 100.0 65.4 70.0 51.6 93.6 60.3 99.9 75.4 93.1 LGrad (2023) 68.6 93.8 69.9 89.2 50.3 54.0 71.1 82.0 57.5 67.3 89.1 99.1 78.5 86.0 78.0 87.4 54.8 68.0 68.6 80.8 Ojha (2023) 78.5 98.3 72.0 98.9 77.6 99.8 77.6 98.9", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1245 }, { "text": "77.6 99.7 78.2 98.7 85.2 98.1 77.6 98.7 74.2 97.8 77.6 98.8 FreqNet 89.8 98.8 98.8 100.0 95.2 98.2 94.5 97.3 95.2 98.2 100.0 100.0 88.3 94.3 85.4 90.5 98.8 100.0 94.0 97.5 Table 2: Cross-model performance on the self-synthesis dataset. Methods Parameters \u2193 mAcc. \u2191of 17 models F3Net(2020) 48.9 M 77.8", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1246 }, { "text": "LGrad(2023) 46.6 M 76.8 Ojha(2023) 304.0 M 83.0 FreqNet 1.9 M 92.8(+9.8) Table 3: Comparison of Parameters. the baselines. We employ the PyTorch framework (Paszke et al. 2019) for the implementation of our method, utiliz- ing the computational power of the Nvidia GeForce RTX 3090 GPU. For the critical task of Fast Fourier Transform", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1247 }, { "text": "(FFT), we leverage the torch.fft.fftn function within the PyTorch library. Deepfake Performance on Real-world Scene In order to demonstrate the remarkable generalization abil- ity of our FreqNet on unseen sources, we carry out evalua- tions on a real-world scene dataset. This dataset comprises images sourced from a total of 17 different generation mod-", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1248 }, { "text": "els, encompassing 8 models from the ForenSynths test set and an additional 9 models from our own self-synthesis pro- (b) StyleGAN (a) ProGAN (c) CelabA-HQ Figure 4: The visualization of Class Activate Map (CAM) (Zhou et al. 2016) extracted from detector on face images. cess. This evaluation setup introduces increased complexity", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1249 }, { "text": "We compare with the previous methods: BiHPF (Jeong et al. 2022a), FreGAN (Jeong et al. 2022c), LGrad (Tan et al. 2023), Ojha (Ojha et al. 2023). To ensure fair and meaning- ful comparisons, we adopt the same experimental setting as the established baselines (Jeong et al. 2022a,c). Specifically, the 1-class, 2-class and 4-class settings refer to training with", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1250 }, { "text": "the (horse), (chair, horse), (car, cat, chair, horse) categories of ProGAN, respectively. The results of the ForenSynths dataset are presented in Table 1. The proposed FreqNet surpasses its counterparts in terms of mean Acc. metric and mean A.P. metrics, except for the value of A.P. on the 1-class setting. Notably, with", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1251 }, { "text": "the 4-class setting, FreqNet achieves a mean Acc. value of 91.5%, demonstrating its strong performance. In compari- son to the current state-of-the-art methods LGrad and Ojha, our FreqNet exhibits substantial improvements, surpassing these methods by 5.4% and 2.4% in mean Acc., using fewer parameters. In the 1- and 2-class settings, our FreqNet also", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1252 }, { "text": "achieves gains of 2.9% and 0.4% compared to Ojha. Fur- thermore, compared to FingerprintNet(Jeong et al. 2022b) tested on six unseen models, our FreqNet achieves a marked improvement in mean accuracy, rising from 82.6% to 90.6% and exhibiting a significant gain of 8.0%. Additionally, we provide results on 9 models from self-synthesis in Table", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1253 }, { "text": "2. We adopt the 4-class setting detector to perform testing. Compared to Ojha, our FreqNet achieves a marked improve- ment in mean accuracy, soaring from 77.6% to an impressive 94.0%, resulting in a significant gain of 16.4%. When test- ing on face images, the proposed FreqNet achieves accuracy rates of 98.7%, 99.0%, and 99.5% on the ProGAN, Style-", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1254 }, { "text": "GAN, and StyleGAN2 datasets, respectively. Furthermore, we provide a comprehensive overview of the number of parameters and the mean accuracy across all real-world scenes in Table 3. It is evident that our Fre- qNet, with a modest parameter count of 1.9 million, signif- icantly outperforms the current state-of-the-art model Ojha", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1255 }, { "text": "(Ojha et al. 2023), which boasts an extensive 304 million parameters. This substantial difference in parameter count translates into a notable performance gain, with our FreqNet achieving a remarkable improvement of 9.8% in mean accu- racy compared to the larger model. This result underscores the efficiency and effectiveness of our FreqNet approach,", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1256 }, { "text": "demonstrating that superior performance can be achieved with significantly fewer parameters, a crucial advantage in real-world applications. Compared to other frequency-based methods, such as BiHPF (Jeong et al. 2022a), FrePGAN (Jeong et al. 2022c), F3Net(Qian et al. 2020), our FreqNet achieves better per-", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1257 }, { "text": "formance on the real-world scene. The results confirm the generalization capability of the proposed frequency domain learning to extract a general representation of artifacts, and generalize this representation across various GAN models and categories. We perform ablation analyses on our FreqNet by individ-", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1258 }, { "text": "ually removing the proposed modules. The results of these ablation experiments are presented in Table 4. Upon removal of the designed modules, we observe a decline in the detec- tion performance, underscoring the efficacy of the proposed components. In the revised version, we will expound further HFRI", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1259 }, { "text": "HFRFS HFRFC FCL mean Acc. \u2713 \u2713 \u2713 84.3 \u2713 \u2713 \u2713 85.3 \u2713 \u2713 \u2713 87.8 \u2713 \u2713 82.0 \u2713 \u2713 \u2713 83.8 \u2713 \u2713 \u2713 \u2713 91.5 Table 4: Ablation Study of FreqNet on the Foren- Synths(Wang et al. 2020). on the specifics of the ablation analysis to provide a more comprehensive understanding. Visualization of Class Activate Map. To visually", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1260 }, { "text": "demonstrate the discriminative regions identified by our de- tector, we present the Class Activation Maps (CAM) in Fig- ure 4. The CAMs are generated using images from ProGAN, StyleGAN, StyleGAN2, and CelebA-HQ. The CAMs pro- vide insights into the areas of focus for the detector in dis- tinguishing real from fake images. It\u2019s noteworthy that the", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1261 }, { "text": "CAMs for real images highlight a broader portion of the image, while the CAMs for fake images tend to emphasize localized regions. Interestingly, even though the detector is primarily trained using a dataset containing cars, cats, chairs, and horses, it showcases the ability to recognize face images", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1262 }, { "text": "effectively. This highlights the versatility and adaptability of our detector in identifying distinct deepfake characteristics, even beyond the classes it was primarily trained on. Conclusion This study has focused on the introduction of FreqNet, a lightweight frequency space learning network designed for", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1263 }, { "text": "the task of generalizable forgery image detection. Our ap- proach capitalizes on the power of frequency domain learn- ing, offering an adaptable solution for the challenging prob- lem of deepfake detection across diverse sources and GAN models. Within our methodology, we introduce practical and compact frequency learning plugin modules designed to", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1264 }, { "text": "compel the CNN classifier to operate within the frequency domain. The extensive experiments conducted on 17 differ- ent generation models serve as compelling evidence of Fre- qNet\u2019s generalization ability. This research contributes to ad- vancing the field of deepfake detection, showcasing the po- tential of FreqNet to effectively combat the challenges posed", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1265 }, { "text": "by evolving forgery techniques and diverse image sources. Acknowledgments This work was supported in part by the National Key R&D Program of China (No.2021ZD0112100), National NSF of China (No.U1936212, No.62120106009), Na- tional Natural Science Foundation of China under Grants 62072394, Natural Science Foundation of Hebei province", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1266 }, { "text": "Fidelity Natural Image Synthesis. In International Confer- ence on Learning Representations. Cao, J.; et al. 2022. End-to-end reconstruction-classification learning for face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 4113\u20134122. Chai, L.; et al. 2020. What makes fake images detectable?", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1267 }, { "text": "understanding properties that generalize. In European con- ference on computer vision, 103\u2013120. Springer. Chen, L.; et al. 2022. Self-supervised learning of adversarial example: Towards good generalizations for deepfake detec- tion. In Proceedings of the IEEE/CVF conference on com- puter vision and pattern recognition, 18710\u201318719.", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1268 }, { "text": "Choi, Y.; et al. 2018. Stargan: Unified generative adversarial networks for multi-domain image-to-image translation. In Proceedings of the IEEE conference on computer vision and pattern recognition, 8789\u20138797. Chollet, F. 2017. Xception: Deep learning with depthwise separable convolutions. In Proceedings of the IEEE con-", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1269 }, { "text": "ference on computer vision and pattern recognition, 1251\u2013 1258. Cooley, J. W.; et al. 1969. The fast Fourier transform and its applications. IEEE Transactions on Education, 12(1): 27\u2013 34. Durall, R.; et al. 2020. Watch your up-convolution: Cnn based generative deep neural networks are failing to repro-", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1270 }, { "text": "duce spectral distributions. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 7890\u20137899. Frank, J.; et al. 2020. Leveraging frequency analysis for deep fake image recognition. In International conference on machine learning, 3247\u20133258. PMLR. Goodfellow, I. J.; et al. 2014. Generative Adversarial Nets.", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1271 }, { "text": "In NIPS. Haliassos, A.; et al. 2021. Lips don\u2019t lie: A generalisable and robust approach to face forgery detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 5039\u20135049. He, Y.; et al. 2021. Beyond the Spectrum: Detecting Deep- fakes via Re-Synthesis. In Proceedings of the Thirtieth", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1272 }, { "text": "International Joint Conference on Artificial Intelligence, IJCAI-21, 2534\u20132541. International Joint Conferences on Artificial Intelligence Organization. He, Z.; et al. 2019. AttGAN: Facial Attribute Editing by Only Changing What You Want. IEEE Transactions on Im- age Processing, 28(11): 5464\u20135478. Jeong, Y.; et al. 2022a.", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1273 }, { "text": "BiHPF: Bilateral High-Pass Fil- ters for Robust Deepfake Detection. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 48\u201357. Jeong, Y.; et al. 2022b. FingerprintNet: Synthesized Finger- prints for Generated Image Detection. In European Confer- ence on Computer Vision, 76\u201394. Springer.", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1274 }, { "text": "Jeong, Y.; et al. 2022c. FrePGAN: robust deepfake detec- tion using frequency-level perturbations. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, 1060\u20131068. Karras, T.; et al. 2018. Progressive Growing of GANs for Improved Quality, Stability, and Variation. In International", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1275 }, { "text": "Conference on Learning Representations. Karras, T.; et al. 2019. A style-based generator architec- ture for generative adversarial networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 4401\u20134410. Karras, T.; et al. 2020. Analyzing and improving the image quality of stylegan. In Proceedings of the IEEE/CVF con-", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1276 }, { "text": "ference on computer vision and pattern recognition, 8110\u2013 8119. Kingma, D. P.; et al. 2015. Adam: A Method for Stochastic Optimization. In ICLR (Poster). Lee, K. S.; et al. 2021. Infomax-gan: Improved adversar- ial image generation via information maximization and con- trastive learning. In Proceedings of the IEEE/CVF winter", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1277 }, { "text": "conference on applications of computer vision, 3942\u20133952. Li, C.; Huang, Z.; Paudel, D. P.; Wang, Y.; Shahbazi, M.; Hong, X.; and Van Gool, L. 2023. A continual deepfake detection benchmark: Dataset, methods, and essentials. In Proceedings of the IEEE/CVF Winter Conference on Appli- cations of Computer Vision, 1339\u20131349.", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1278 }, { "text": "Li, C.-L.; et al. 2017. Mmd gan: Towards deeper under- standing of moment matching network. Advances in neural information processing systems, 30. Li, J.; et al. 2021. Frequency-aware discriminative feature learning supervised by single-center loss for face forgery detection. In Proceedings of the IEEE/CVF conference on", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1279 }, { "text": "computer vision and pattern recognition, 6458\u20136467. Li, Y.; et al. 2018. In ictu oculi: Exposing ai created fake videos by detecting eye blinking. In 2018 IEEE Inter- national workshop on information forensics and security (WIFS), 1\u20137. IEEE. Lin, T.-Y.; et al. 2014. Microsoft coco: Common objects in", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1280 }, { "text": "context. In European conference on computer vision, 740\u2013 755. Springer. Liu, M.; et al. 2019. Stgan: A unified selective transfer net- work for arbitrary image attribute editing. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 3673\u20133682. Liu, Z.; et al. 2015. Deep learning face attributes in the", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1281 }, { "text": "Lu\u02c7ci\u00b4c, M.; et al. 2019. High-fidelity image generation with fewer labels. In International conference on machine learn- ing, 4183\u20134192. PMLR. Luo, Y.; et al. 2021. Generalizing face forgery detection with high-frequency features. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition,", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1282 }, { "text": "16317\u201316326. Masi, I.; et al. 2020. Two-branch recurrent network for iso- lating deepfakes in videos. In European conference on com- puter vision, 667\u2013684. Springer. Miyato, T.; et al. 2018. Spectral normalization for generative adversarial networks. arXiv preprint arXiv:1802.05957. Nie, W.; et al. 2019. Relgan: Relational generative adversar-", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1283 }, { "text": "ial networks for text generation. In International conference on learning representations. Ojha, U.; et al. 2023. Towards universal fake image detectors that generalize across generative models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 24480\u201324489. Park, T.; et al. 2019.", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1284 }, { "text": "Semantic image synthesis with spatially-adaptive normalization. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2337\u20132346. Paszke, A.; et al. 2019. Pytorch: An imperative style, high- performance deep learning library. Advances in neural in- formation processing systems, 32.", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1285 }, { "text": "Qian, Y.; et al. 2020. Thinking in frequency: Face forgery detection by mining frequency-aware clues. In European conference on computer vision, 86\u2013103. Springer. Rossler, A.; et al. 2019. Faceforensics++: Learning to detect manipulated facial images. In Proceedings of the IEEE/CVF international conference on computer vision, 1\u201311.", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1286 }, { "text": "Russakovsky, O.; et al. 2015. Imagenet large scale visual recognition challenge. International journal of computer vi- sion, 115(3): 211\u2013252. Shiohara, K.; et al. 2022. Detecting deepfakes with self- blended images. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, 18720\u2013", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1287 }, { "text": "18729. Tan, C.; et al. 2023. Learning on Gradients: Generalized Artifacts Representation for GAN-Generated Images De- tection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 12105\u2013 12114. Wang, C.; et al. 2021. Representative forgery mining for fake face detection. In Proceedings of the IEEE/CVF con-", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1288 }, { "text": "ference on computer vision and pattern recognition, 14923\u2013 14932. Wang, S.-Y.; et al. 2020. CNN-generated images are surprisingly easy to spot... for now. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 8695\u20138704. Wang, Y.; et al. 2023. Dynamic Graph Learning With", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1289 }, { "text": "Content-Guided Spatial-Frequency Relation Reasoning for Deepfake Detection. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, 7278\u2013 7287. Woo, S.; et al. 2022. ADD: Frequency Attention and Multi- View Based Knowledge Distillation to Detect Low-Quality Compressed Deepfake Images. In Proceedings of the AAAI", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1290 }, { "text": "Conference on Artificial Intelligence, volume 36, 122\u2013130. Yu, F.; et al. 2015. Lsun: Construction of a large-scale image dataset using deep learning with humans in the loop. arXiv preprint arXiv:1506.03365. Yu, Y.; et al. 2020. Mining generalized features for detecting ai-manipulated fake faces. arXiv preprint", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1291 }, { "text": "arXiv:2010.14129. Zhou, B.; et al. 2016. Learning deep features for discrimi- native localization. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2921\u20132929. Zhu, J.-Y.; et al. 2017. Unpaired image-to-image translation using cycle-consistent adversarial networks. In Proceedings", "source": "FreqNet Frequency-Aware Detection", "year": 2024, "url": "https://arxiv.org/abs/2403.07240", "id": 1292 }, { "text": "Deepfake Detection: A Comprehensive Survey from the Reliability Perspective TIANYI WANG, Department of Computer Science, The University of Hong Kong, China XIN LIAO, College of Computer Science and Electronic Engineering, Hunan University, China KAM PUI CHOW, Department of Computer Science, The University of Hong Kong, China", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1293 }, { "text": "XIAODONG LIN, School of Computer Science, University of Guelph, Canada YINGLONG WANG, Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Qilu University of Technology (Shandong Academy of Sciences), China The mushroomed Deepfake synthetic materials circulated on the internet have raised a profound social impact", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1294 }, { "text": "on politicians, celebrities, and individuals worldwide. In this survey, we provide a thorough review of the existing Deepfake detection studies from the reliability perspective. We identify three reliability-oriented research challenges in the current Deepfake detection domain: transferability, interpretability, and robustness.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1295 }, { "text": "Moreover, while solutions have been frequently addressed regarding the three challenges, the general reliability of a detection model has been barely considered, leading to the lack of reliable evidence in real-life usages and even for prosecutions on Deepfake-related cases in court. We, therefore, introduce a model reliability study", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1296 }, { "text": "metric using statistical random sampling knowledge and the publicly available benchmark datasets to review the reliability of the existing detection models on arbitrary Deepfake candidate suspects. Case studies are further executed to justify the real-life Deepfake cases including different groups of victims with the help of", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1297 }, { "text": "the reliably qualified detection models as reviewed in this survey. Reviews and experiments on the existing approaches provide informative discussions and future research directions for Deepfake detection. CCS Concepts: \u2022 Security and privacy \u2192Human and societal aspects of security and privacy; \u2022 Computing methodologies \u2192Computer vision; \u2022 Applied computing \u2192Computer forensics; \u2022", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1298 }, { "text": "General and reference \u2192Surveys and overviews. Additional Key Words and Phrases: Deepfake detection, reliability study, forensic investigation, confidence interval 1 INTRODUCTION In June 2022, a mother from Pennsylvania, known as a \u2018Deepfake mom\u2019, is sentenced to probation for three years because of her harassment of the rivals on her daughter\u2019s cheerleader team [87].", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1299 }, { "text": "She is originally accused of using the so-called Deepfake technology to generate and spread fake videos of her daughter\u2019s opponents depicting indelicate behaviors in March 2021. However, as admitted by the prosecutors, the justification for generating Deepfake videos is unable to be confirmed without accurate evidence and tools [62]. The term Deepfake refers to a deep learning", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1300 }, { "text": "technique raised by the Reddit user \u2018deepfakes\u2019 [34] in 2017 [143] that could automatically execute face-swapping from a source person to a target one while maintaining all other contents of the target image unchanged including expression, movement, and background scene. Later, the face reenactment technology [84, 169, 183, 191], which transfers attributes of a source face to a target one", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1301 }, { "text": "while maintaining the target\u2019s facial identity, is also classified as Deepfake from a comprehensive standpoint. Corresponding authors: Yinglong Wang and Xin Liao. Authors\u2019 addresses: Tianyi Wang, terry.ai.wang@gmail.com, Department of Computer Science, The University of Hong Kong, Pok Fu Lam, Hong Kong, China; Xin Liao, xinliao@hnu.edu.cn, College of Computer Science and Electronic Engineering,", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1302 }, { "text": "Hunan University, Address, Changsha, Hunan, China, 410082; Kam Pui Chow, chow@cs.hku.hk, Department of Computer Science, The University of Hong Kong, Pok Fu Lam, Hong Kong, China; Xiaodong Lin, xlin08@uoguelph.ca, School of Computer Science, University of Guelph, 50 Stone Road East, Guelph, Ontario, Canada, N1G 2W1; Yinglong Wang,", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1303 }, { "text": "2 Wang et al. Deepfake can bring benefits and convenience to people\u2019s daily lives, especially from the per- spective of human entertainment. Specifically, movie lovers may swap their faces onto movie clips and perform as their favorite superheroes and superheroines [89]. On the other hand, a tainted", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1304 }, { "text": "celebrity who is no longer allowed to appear on TV shows [161\u2013163] can be face-swapped within completed TV productions instead of reshooting or removing the episode. Moreover, Deepfake is available for bringing a deceased person digitally back to life [5] and greeting families and friends with desired conversations. Besides, e-commerce is another scenario to exert positive effects of", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1305 }, { "text": "Deepfake by trying on clothes in an online fitting room [178]. Benefiting from the publicly available source code implementations on the internet, various Deep- fake mobile applications [76, 96, 110] have been released. The early product in 2018, FakeApp [33], requires a large number of input images to achieve satisfactory synthetic results. Later in 2019, the", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1306 }, { "text": "popular Chinese application, ZAO [122], can generate face-swapping outputs by simply inputting a series of selfies. Recently, a more powerful facial synthetic tool, Reface [152], is built with further functionalities such as moving and singing animations based on hyper-realistic synthetic results. Although entertaining human lives, the free-access Deepfake applications require practically no", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1307 }, { "text": "experience in the field when generating fake faces, which therefore poses crucial potential threats to society. Because of the hyper-realistic quality that is indistinguishable by human eyes [126\u2013128], Deepfake has already been ranked as the most serious artificial intelligence crime threat since 2020 [148], and the current and potential victims include politicians, celebrities, and even every", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1308 }, { "text": "human being on earth. Besides the experimental fake Obama1, a fake president Zelensky [120] has caused panic in Ukraine during the Russia-Ukraine war. As for celebrities, fake porn videos have frequently targeted female actresses and representative victims include Emma Watson, Natalie Portman, and Ariana Grande [88, 97]. Additionally, remember the \u2018Deepfake mom\u2019 as discussed at", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1309 }, { "text": "the beginning? It is one of the best proofs that anyone can become a victim of Deepfake in modern society. To protect individuals and society from the negative impacts of misusing Deepfake, Deepfake detection approaches have been frequently designed and they mostly conduct a binary classification task to identify real and fake faces with the help of deep neural networks (DNNs). In this survey, we", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1310 }, { "text": "provide an in-depth review of the developing Deepfake detection approaches from the reliability perspective. In other words, we focus on studies and topics that devote to the ultimate success of Deepfake detection in real-life usages, and more importantly, for criminal investigations in court-case judgments. Early studies mainly focus on the in-dataset model performance such that", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1311 }, { "text": "the detection models are trained and validated the performance on the same dataset. While most recent work has achieved promising detection performance for the in-dataset test, the research challenges at the current stage for Deepfake detection can be concluded in three aspects, namely, transferability, interpretability, and robustness. The transferability topic refers to the progress of", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1312 }, { "text": "improving the cross-dataset ability of models when evaluated on unseen data. As the detection performance keeps advancing by various approaches, interpretability is another research goal to explain the reason that the detection model determines the falsification. Moreover, when applying well-trained and well-performed detection models for real-life scenarios, robustness is considered a", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1313 }, { "text": "main topic in dealing with various real-life conditions. While research has been incrementally attempted regarding the three challenging topics, a reliable Deepfake detection model is expected to benefit people\u2019s daily lives with good transferabil- ity on unseen data, have convincing interpretability of the detection decision, and show robust", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1314 }, { "text": "Deepfake Detection: A Comprehensive Survey from the Reliability Perspective 3 without an authenticated scheme to nominate the detection models as reliable evidence to assist prosecutions and judgments in court, similar failures as the \u2018Deepfake mom\u2019 case will happen again due to the lack of reliable detection tools to support the accusation of Deepfake even though", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1315 }, { "text": "detection performance on each benchmark dataset is reported in current work. In other words, the trustworthiness of a model-derived falsification needs to be proved before it can convince people in real-life usages and for court-case judgments. To fulfill the research gap of the model reliability study, beyond the comprehensive review of Deepfake detection, we devise a scheme to scientifically", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1316 }, { "text": "validate the reliability of the well-developed Deepfake detection models using statistical random sampling knowledge [117]. To guarantee the credibility of the reliability study, we concurrently introduce a systematic workflow of data pre-processing including image frame selection and extrac- tion from videos and face detection and cropping, which has been barely mentioned with concrete", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1317 }, { "text": "details in past work. Thereafter, we quantitatively evaluate and record the selected state-of-the-art Deepfake detection approaches by training and testing with their reported optimal settings on the same group of pre-processed datasets in a completely fair game. Thenceforth, we validate the detection model reliability following the designed evaluation scheme. In the end, a case study is", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1318 }, { "text": "enforced to justify the results derived by the detection models on four well-known real-life synthetic videos concerning the reliable detection accuracies statistically at 90% and 95% confidence levels based on the research outcomes from the model reliability study. Interesting findings and future research topics that have been scarcely concluded in previous studies [42, 89, 121, 164, 165, 178]", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1319 }, { "text": "are analyzed and discussed. Furthermore, we believe that the proposed Deepfake detection model reliability study scheme is informative and can be adopted as evidence to assist prosecutions in court once granted approval by authentication experts or institutions following the legislation. The rest of the paper is organized as follows. We provide a brief review of the popular synthetic", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1320 }, { "text": "techniques and publicly available benchmark datasets in Section 2. Then, we define the challenges of the Deepfake detection research and provide a thorough review of the development history of the Deepfake detection approaches in Section 3. In Section 4, we illustrate the model reliability study scheme and demonstrate the algorithm details. In Section 5, we detailedly introduce a standardized", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1321 }, { "text": "data pre-processing workflow and list the participating datasets in the experiments of this paper. In Section 6, we conduct detection performance evaluation and reliability justification using selected state-of-the-art models on the benchmark datasets and following the reliability study scheme, respectively. Section 7 exhibits Deepfake detection results of the selected models when applying", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1322 }, { "text": "to the real-life videos in a case study, and discussions along with experiment results from early sections are presented in Section 8. Section 9 concludes the remarks and highlights the potential future directions in the research domain. 2 THE EVOLUTION OF DEEPFAKE GENERATION 2.1 Deepfake Generation", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1323 }, { "text": "Deepfake is initially raised in the Reddit community when the open-source implementation was first published by the user \u2018deepfakes\u2019 simultaneously in 2017. Early research mainly focuses on subject-specific approaches, which can only swap facial identities that the models have seen during training. The most popular framework of the existing Deepfake synthesis studies [34, 137] for the", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1324 }, { "text": "subject-specific identity swap is an autoencoder [91]. In a nutshell, the autoencoder contains a shared encoder that extracts identity-independent features from the source and target faces and two unique decoders each is in charge of reconstructing synthetic faces of a corresponding facial identity. Specifically, in the training phase, faces of the source identity are fed to the encoder for", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1325 }, { "text": "4 Wang et al. face reconstruction is trained following the same workflow. When using a well-trained model to operate face-swapping, a target face after context vector extraction is fed to the decoder that reconstructs the source identity. Thenceforth, the decoder generates a look maintaining the facial", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1326 }, { "text": "expression and movement of the target face while having the identity of the desired source face. If the face-swapping model is trained for both directions, a target face may be face-swapped onto a source face following the same workflow. Recent studies gradually focus on subject-agnostic methods to enable face-swapping for arbitrary", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1327 }, { "text": "identities with higher resolutions and have exploited Generative Adversarial Networks (GAN) [50] for better synthesis authenticity [10, 23, 36, 49, 85, 86, 123, 129, 134, 188, 195]. In other words, they aim to consistently produce high-quality face-swapping results even on facial identities that are unseen during model training. GAN is a generator-discriminator architecture that is trained", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1328 }, { "text": "by having two components battle against each other to advance the output quality. In practice, the generator is periodically trained to fool the discriminator with synthetic faces. For instance, FaceShifter [100] and SimSwap [19] each devise particular modules to preserve facial attributes that are hard to reconstruct and maintain fidelity for arbitrary facial identities. MegaFS [196] and", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1329 }, { "text": "HifiFace [176] accomplish face-swapping at high resolutions of 512 and 1,024 for arbitrary facial identities, respectively, relying on the reconstruction ability of GAN. In the last step, the generated fake face is usually blended back to the pristine target image with tuning techniques such as blurring and smoothing [190] to reduce the visible Deepfake traces.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1330 }, { "text": "2.2 Deepfake Benchmark Datasets Benchmark datasets are vital in the development history of Deepfake detection models. Dolhansky et al. [197] raised the idea to break down previous datasets into three generations based on the two-generation categorization in the early work [81]. As listed in Table 1, UADFV [186], Deepfake-", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1331 }, { "text": "TIMIT [93], and FaceForensics++ (FF++) [144] are categorized to the first generation; Deepfake Detection Dataset [57], DeepFake Detection Challenge (DFDC) Preview [38], and Celeb-DF [104] are in the second generation; DeeperForensics-1.0 (DF1.0) [81] and DeepFake Detection Chal- lenge (DFDC) [37] are in the third generation. In summary, later generations contain general", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1332 }, { "text": "improvements over the previous ones in terms of dataset magnitude or synthetic method diversity. Unlike the summary by Dolhansky et al. [197], the agreement from individuals appearing is not considered in this study, and we instead re-define the third generation such that the datasets are of better quality, broader diversity and magnitude, higher difficulty than the early generations, or", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1333 }, { "text": "challenging detailed discrepancies in early synthetic videos are resolved in the datasets. Specifically, besides DFDC with large manipulation diversity and dataset magnitude and DF1.0 with large magnitude and considerable difficulty by adding deliberate perturbations, we further classify the following datasets in the third generation, namely, FaceShifter [100], WildDeepfake [197], and", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1334 }, { "text": "KoDF [95]. FaceShifter [100], although synthesized based on real videos of FF++, has specifically solved the so-called facial occlusion challenge that appears in previous datasets. In other words, the synthetic results are better handled even in difficult cases where parts of the face are blocked or obscured by", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1335 }, { "text": "objects such as accessories or other body parts such as hair tippings. WildDeepfake (WDF) [197] appears to be a special one in the third generation because videos are totally collected from the internet, which matches the real-life Deepfake circumstance the best. The most recent KoDF [95] dataset is so far the largest Deepfake benchmark dataset that is publicly available with reasonable", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1336 }, { "text": "Deepfake Detection: A Comprehensive Survey from the Reliability Perspective 5 Table 1. Information of the existing Deepfake datasets categorized into three generations based on quality, diversity, and difficulty. Publication year, the number of real and fake sequences, and the source of real and fake materials are listed.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1337 }, { "text": "Dataset Year # Real / Fake Real / Fake Source Generation UADFV [186] 2018 49 / 49 YouTube / FakeApp [33] 1st Generation DeepfakeTIMIT [93] 2018 \u2013 / 620 faceswap-GAN [111] FF++ [144] 2019 1,000 / 4,000 YouTube / 4 methods2 DFD [57] 2019 363 / 3,068 consenting actors / unknown methods 2nd Generation DFDC Preview [38]", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1338 }, { "text": "2019 1,131 / 4,119 crowdsourcing / 2 unknown methods Celeb-DF [104] 2019 590 / 5,639 YouTube / improved Deepfake DF1.0 [81] 2020 \u2013 / 10,000 FF++ real / DF-VAE 3rd Generation FaceShifter [100] 2020 \u2013 / 1,000 FF++ real / GAN-based DFDC [37] 2020 23,654 / 104,500 crowdsourcing / 8 methods3 WDF [197] 2020", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1339 }, { "text": "3,805 / 3,509 video-sharing websites KoDF [95] 2021 62,166 / 175,776 lab-controlled / 6 manipulations4 3 RELIABILITY-ORIENTED CHALLENGES OF DEEPFAKE DETECTION Detection work on Deepfake has been proposed since the first occurrence of Deepfake contents. Classical forgery detection approaches [28, 35, 43, 47, 133, 135] mainly focus on the intrinsic", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1340 }, { "text": "statistics and hand-crafted traces such as eye blinking [82, 102], head pose [186], and visual artifacts [119] to analyze the spatial feature manipulation patterns. Besides, there are papers that have derived high accuracies and AUC scores by training and testing on the same dataset of a synthetic method. Several studies [60, 145] integrated CNN and Long Short-term Memory", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1341 }, { "text": "(LSTM) [68] for spatial and temporal features analyses, respectively, and accomplished detection for in-dataset evaluations on self-collected data and the FF++ dataset. Hsu, Zhang, and Lee [70] utilized GAN-generated fake samples for real-fake pairwise training using DenseNet [74]. Agarwal et al. [3]", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1342 }, { "text": "accomplished detection on face-swap Deepfake using CNNs with biometric information including facial expressions and head movements. However, although accomplished well-pleasing in-dataset detection performance on some early or self-collected datasets, they are mostly easily fooled by the hyper-realistic Deepfake contents in the current research domain because of limitations in dataset", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1343 }, { "text": "quality, dataset diversity, and method or model ability. Later studies gradually consider Deepfake detection as a binary classification task using DNNs. As malicious Deepfake contents have started to jeopardize human society and the cases are even discussed in court, reliably trusted detection methods are eagerly desired by the public. In particular,", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1344 }, { "text": "three challenges (Fig. 1) of the current Deepfake detection research domain can be summarized regarding the reliability goal, namely, transferability, interpretability, and robustness. 3.1 Transferability Deep learning models usually exhibit satisfactory performance on the same type of data that are", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1345 }, { "text": "seen in the training process but perform poorly on unseen data. In real life, Deepfake materials can be generated via various synthetic techniques [19, 49, 111, 115, 196] as abundant on-the-shelf 2FaceSwap, Deepfakes, Face2Face, and NeuralTextures. 3DF-128, DF-256, MM/NN, NTH, FSGAN, StyleGan, refinement, and audio swaps.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1346 }, { "text": "6 Wang et al. Fig. 1. Demonstrations of the three challenges from top to bottom. Transferability (top) refers to models that focus on stable detection ability on unseen benchmark datasets; interpretability (middle) refers to efforts on explaining the model detected falsification; robustness (bottom) refers to models that handle Deepfake", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1347 }, { "text": "suspects under different real-life conditions and scenarios. easily accessible face-swapping implementations are publicly available, and a reliable Deepfake detection model is expected to perform well on unseen data to imitate real-life Deepfake cases. Therefore, guaranteeing the transferability of the detection models for cross-dataset performance", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1348 }, { "text": "is necessary and frequently discussed. With the fast development of deep learning techniques, various methods are devised using simple Convolutional Neural Network (CNN) based models. Zhou et al. [194] fused a CNN stream with a support vector machine (SVM) [66] stream to analyze face features with the assistance of", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1349 }, { "text": "local noise residuals. Afchar et al. [2] studied the mesoscopic features of images with successive convolutional and pooling layers. Nguyen et al. [124] designed a multi-task learning scheme to simultaneously perform detection and localization using CNNs. DFT-MF [77] thinks that the mouth features can be important for detection and utilizes a convolutional model to detect Deepfake by", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1350 }, { "text": "Deepfake Detection: A Comprehensive Survey from the Reliability Perspective 7 x-ray on the candidate Deepfake faces by revealing whether the blending of two different source images can be decomposed. Rather than the basic CNN architectures, well-designed and pre-trained CNN backbones are frequently exploited to improve model performance on Deepfake detection, especially for the", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1351 }, { "text": "cross-dataset performance on unseen data as more benchmark datasets have been released. An early approach [4] proposes optical flow analysis with pre-trained VGG16 [153] and ResNet50 [64] CNN backbones and achieves preliminary in-dataset test performance on the FaceForensics++ (FF++) dataset [144]. Capsule [125] employs capsule architectures [146] with light VGG19-based", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1352 }, { "text": "network parameters but achieves similar detection performance to the traditional approaches leveraging CNNs. Li et al. [103] conducted a strengthened model, DSP-FWA, with the help of the spatial pyramid pooling [63] with ResNet50 as the backbone. This method is shown to be applicable to Deepfake materials at different resolution levels. FFD [31] leverages the popular", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1353 }, { "text": "attention mechanism by element-wise multiplication to study the feature maps and utilizes the XceptionNet CNN backbone, achieving marginally more promising performance than the work by Rossler et al. [144]. SSTNet [184] exploits Spatial, Steganalysis, and Temporal features using XceptionNet [24] and exhibits reasonable intra-cross-dataset performance on FF++. Given the", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1354 }, { "text": "assumption that Deepfake only modifies the internal part of the face, Nirkin et al. [130] adopted two streams of XceptionNet [24] for face and context (hair, ears, neck) identification and another XceptionNet to classify real and fake based on the learned discrepancies between the two. Later, Bonettini et al. [9] and Tariq et al. [158] studied the ensemble of various pre-trained convolutional", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1355 }, { "text": "models including EfficientNetB4 [157], XceptionNet [24], DenseNet [74], VGGNet [153], ResNet [64], and NASNet [198] backbones to detect Deepfake. Rossler et al. [144] employed the pre-trained well-designed XceptionNet [24] network and achieved state-of-the-art detection performance at the time on FF++. Besides, a special study introduced by Wang et al. [171] proves the transferability", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1356 }, { "text": "of the model trained on one CNN-generated dataset to the rest ten using pre-trained ResNet50 [64]. Meanwhile, frequency cues are also noticed and analyzed by researchers. While early image forgery detection work [6] focus on all high, medium, and low frequencies via Fourier transform, Frank et al. [46] were the first that raised the idea of finding frequency inconsistency between real", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1357 }, { "text": "and fake by employing low-frequency features to assist Deepfake detection using RGB information. Since low-frequency traces are mostly hidden by blurry facial features, later studies mainly analyze high-level features. F3-Net [141] exploits both low- and high-frequency features without RGB feature extraction operation. DFT [41], Two-Branch [118], SPSL [107], and MPSM-RFAM [20]", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1358 }, { "text": "accomplish promising detection results by analyzing high-frequency spectrum along with the on-the-shelf RGB features. Li et al. [99] extracted both middle- and high-frequency features and correlated frequency and RGB features for detection. By pointing out the drawbacks of using coarse-grained frequency information, Gu et al. [53] combined fine-grained frequency information", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1359 }, { "text": "with RGB features for better feature richness in a latest work. Moreover, Jeong et al. [79] designed a new training scheme with frequency-level perturbation maps added, which further enhanced the generalization ability of the detection model regarding all GAN-based generators. Since the CNN architecture and backbones lack generalization ability and mainly focus on local", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1360 }, { "text": "features, even XceptionNet is restricted in learning the global features for further performance improvements. Therefore, developed solutions have introduced convolutional spatial attention to enlarge the local feature area and learn the corresponding relations, and the detection AUC scores have been gradually raised above 70% on average upon unseen datasets accordingly. The SRM [113]", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1361 }, { "text": "approach makes two streams of network using XceptionNet as the CNN backbone and focuses on the high-frequency information and RGB frames with a spatial attention module in each stream. The MAT model [192] proposes the CNN backbone EfficientNetB4 and borrows the convolutional attention idea to study different local patches of the input image frame. Specifically, the artifacts", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1362 }, { "text": "8 Wang et al. in shallow features are zoomed in for fine-grained enhancement in the detection performance. With the success of the transformer architecture [170] in the natural language processing (NLP) domain, different versions of vision transformer [40, 108, 136, 167, 175] have derived reasonable", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1363 }, { "text": "performance in the computer vision domain due to its ability on global image feature learning. The architectures of vision transformers (ViT) have also been employed for Deepfake detection with promising results in recent work [21, 67, 78, 172, 179]. While reaching the bottleneck specially for the cross-dataset performance even resorting to the", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1364 }, { "text": "advanced powerful but potentially time-consuming neural networks, most recent work gradually focuses on strategies to enrich diversity in training data. Such attempts have improved the detection AUC scores even up to 80% on some unseen datasets. PCL [193] is introduced with an inconsistency image generator to add synthetic diversity and provide richly annotated training data. Sun et", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1365 }, { "text": "al. [155] proposed Dual Contrastive Learning (DCL) to study positive and negative paired data and enrich data views for Deepfake detection with better transferability. FInfer [71] inferences future image frames within a video and is trained based on the representation-prediction loss. Shiohara and Yamasaki [151] presented a novel synthetic training data called self-blended images", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1366 }, { "text": "(SBIs) by blending solely pristine images from the original training dataset, and classifiers are thus boosted to learn more generic representations. Chen et al. [17] proposed the SLADD model to further synthesize the original forgery data for model training to enrich model transferability on unseen data using pairs of pristine images and randomly selected forgery references with forgery", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1367 }, { "text": "configurations including forgery region, blending type, and mix-up blending ratio. Cao et al. [13] introduced the RECCE model to emphasize the common compact representations of genuine faces based on reconstruction-classification learning on real faces. The reconstruction learning over real images enhances the learning representations to be aware of unknown forgery patterns. Liang et", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1368 }, { "text": "al. [105] proposed an easily embeddable disentanglement framework to remove content information while maintaining artifact information for training and Deepfake detection using reconstructed data with various combinations of content and artifact features of real and fake samples. The OST [18] method improves the detection performance by preparing pseudo-training samples based", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1369 }, { "text": "on testing images to update the model weights before applying to the true samples. 3.2 Interpretability Despite the promising ability of deep learning models, they suffer the weak interpretability problems due to the black-box characteristic [109]. In other words, it is hard to explain how and why a model", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1370 }, { "text": "comes up with a particular result. For the Deepfake detection task, while achieving promising detection performance statistically for in- and cross-dataset evaluations, the interpretability issue is still maintained to be fully resolved at the current stage. In other words, people tend to trust methods that are easily understandable via common sense rather than those with satisfactory", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1371 }, { "text": "accuracies but derived based on features that are hard to explain. Consequently, forensic evidence is critical to be probed to interpret and support the detection model performance by highlighting the reasons for classifying fake samples. In a nutshell, the interpretability challenge is to answer the following questions regarding a Deepfake suspect in order to be reliable:", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1372 }, { "text": "\u2022 Why is the content classified as fake? \u2022 Based on which part does the detection model determine the content as fake? As confirmed by Baldassarre et al. [7] using a series of quantitative metrics for evaluating the interpretability of the detection models, heatmaps are generally futile and unacceptable when being", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1373 }, { "text": "Deepfake Detection: A Comprehensive Survey from the Reliability Perspective 9 traces for explanations upon the detection results. Early work mostly relies on the Photo Response Non-Uniformity (PRNU), a small factory-defects-generated noise pattern in the light-sensitive sensors of a digital camera [112]. PRNU has shown strong abilities in source anonymization [138]", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1374 }, { "text": "and source device identification [116, 147]. Unfortunately, most of the PRNU-based Deepfake detection studies [32, 92] have failed to show strong detection performance statistically. Therefore, the PRNU noise pattern can be a useful instrument for source device identification tasks, but it may not be a meaningful forensic noise tracing tool to satisfy the purpose of the Deepfake detection", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1375 }, { "text": "task with respect to the interpretability goal. Later approaches [4, 16, 77, 101] step up to utilize CNNs for noise extraction and analyze the trace differences between real and fake. A milestone denoiser, DnCNN, proposed by Zhang et al. [189] is able to perform blind Gaussian denoising with promising performance, and has been", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1376 }, { "text": "later extended to study the camera model fingerprint for various downstream tasks including forgery detection and localization for face-swapping images based on the underlying noises [29, 30]. Recently, studies [55, 56] aim to extract the manipulation traces for detection and interpretation. Guo et al. [58] proposed the AMTEN method to suppress image content and highlight manipulation", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1377 }, { "text": "traces as more filter iterations are applied. Wang et al. [173, 174] utilized the pre-trained denoisers for Deepfake forensic noise extraction and investigated the underlying Deepfake noise pattern consistency between face and background squares with the help of the siamese architecture [11]. Guo et al. [59] proposed a guided residual network to maintain the manipulation traces within the", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1378 }, { "text": "guided residual image and analyzed the difference between real and fake. Besides studies [26, 140] that rely on biological signals by analyzing minuscule periodic changes through the faces visualizing distinguishable sequential signal patterns as indicators, other studies mostly attempt to explore universal and representative artifacts directly from the visual concepts", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1379 }, { "text": "of the Deepfake materials. The PRRNet [150] approach studies region-wise relation and pixel-wise relation within the candidate image for detection and can roughly locate the manipulated region using pixel-wise values. Trinh et al. [168] fed the dynamic prototype to the detection model and successfully visualized obvious artifact fluctuation via the prototype. Yang et al. [65] proposed", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1380 }, { "text": "to re-synthesis testing images by incorporating a series of visual tasks using GAN models, and they finally extracted visual cues to help perform detection. Obvious artifact differences can be observed by their Stage5 model on real and fake samples. Most recently, Dong et al. [39] introduced FST-matching to disentangle source-, target-, and artifact-relevant features from the input image", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1381 }, { "text": "and improved the detection performance by utilizing artifact-relevant features solely. The adopted features within the fake samples are visualized to explain the detection results. 3.3 Robustness As the transferability and interpretability challenges have been frequently undertaken and reason- able or even promising results have been achieved accordingly, there lacks a stage to be fulfilled", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1382 }, { "text": "in order for the detection approaches to be useful in real-life cases. The challenge of this stage can be summarized as robustness. To be specific, the quality of the candidate Deepfake mate- rial in real life is not as ideal as the benchmark datasets in experimental conditions most of the time. Consequently, they may experience different levels of compression operation in multiple", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1383 }, { "text": "scenarios due to objectively limited conditions [114]. Moreover, post-processing strategies and artificially added perturbations can cause further challenges and even disable the well-trained and well-performed detection models. Therefore, the robustness of the detection approaches is necessary to be significantly considered when facing various real-life application scenarios under", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1384 }, { "text": "10 Wang et al. limitations such as video compression due to network flow settings. In an early study, Kumar, Vatsa, and Singh [94] designed multiple streams of ResNet18 [64] to specifically deal with face reenactment manipulation at various compression levels. Hu et al. [72] proposed a two-stream method that specifically analyzes the features of compressed videos that are widely spread on social", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1385 }, { "text": "networks. The LRNet [156] model is developed to stay robust when detecting highly compressed or noise-corrupted videos. Cao et al. [14] solved the difficulty of detecting against compressed videos by feeding compression-insensitive embedding feature spaces that utilize raw and compressed forgeries to the detection model. Wu et al. [181, 182] analyzed noise patterns generated via online", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1386 }, { "text": "social networks before feeding image data into the detection model for training. The method has won top ranking against existing approaches especially when facing forgeries after being transmitted through social networks. Le and Woo [8] employed attention distillation in the frequency perspective and have successfully raised the detection ability on highly compressed data using ResNet50.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1387 }, { "text": "RealForensics [61] aims to solve Deepfake contents with real-life quality by training the detection model with an auxiliary dataset containing real talking faces before utilizing the benchmark datasets. The method has significantly advanced the detection performance against multiple objective scenarios and adversarial attacks such as Gaussian noise and video compression. Besides, a recent", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1388 }, { "text": "work [177] has constructed a new dataset containing Deepfake contents under the near-infrared condition to prevent potential future Deepfake attacks in the corresponding scenarios. On the other hand, the active condition summarizes deliberate adversarial attacks such as distortions and perturbations. Gandhi and Jain [48] explored Lipschitz regularization [180] and", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1389 }, { "text": "Deep Image Prior (DIP) [98] to remove the artificial perturbations they generated and maintain model robustness. Yang et al. [25] simulated commonly seen data corruption techniques on the benchmark datasets to increase data diversity in model training. Operations such as resolution down-scaling, bit-rate adjustment, and artifacting have effectively boosted the detection ability", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1390 }, { "text": "against in-the-wild Deepfake contents. Hooda et al. [69] designed a Disjoint Deepfake Detection (D3) detector that improves adversarial robustness against artificial perturbations using an ensemble of models. The LTTD [54] framework is enforced to conquer the challenges brought by post-processing procedures of Deepfake generation like visual compression to models that rely on low-level image", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1391 }, { "text": "feature patterns. Lin et al. [106] proved the fact that having temporal information to participate in detection makes the detector less prone to black-box attacks. Moreover, in the latest studies, researchers have frequently illustrated the necessity of robustness by fooling the well-trained detection models with stronger adversarial attacks [51, 75]. Jia et al. [80]", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1392 }, { "text": "even emphasized the robustness of the detection models against potential adversarial attacks by injecting frequency perturbations to fool the state-of-the-art approaches. Also, the robustness of Deepfake detectors is evaluated via the Fast Gradient Sign Method (FGSM) and the Carlini- Wagner L2-norm attack in several studies [48, 149]. Moreover, Carlini and Farid [15] employed", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1393 }, { "text": "white- and black-box attacks on a well-trained detector in five case studies. Obvious performance damping can be observed from the reported results. Faces with imperceptible visual variation even after perturbations and noises are added in have highlighted the importance of studying model robustness in the current Deepfake detection research domain.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1394 }, { "text": "4 DETECTION MODEL RELIABILITY STUDY 4.1 Overview In the ideal scenario, a reliable Deepfake detection model should retain promising transferability on unseen data with unknown synthetic techniques, pellucid explanation upon the falsification, and robust resistance against different real-life application conditions. As reviewed in Section 3, despite", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1395 }, { "text": "Deepfake Detection: A Comprehensive Survey from the Reliability Perspective 11 a method favorably satisfying all three challenges simultaneously is not yet accomplished, a metric that nominates the reliability of a method for real-life usages and court-case judgments is needed. Regardless of the largely improved but still unsatisfied cross-dataset performance in the evolution", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1396 }, { "text": "of Deepfake detection, existing studies have only evaluated the model performance on each testing dataset to show the model detection abilities while the values for each evaluation metric (accuracy and AUC score) vary depending on different testing sets adopted in experiments. On the contrary, in real-life cases, people have no clue about fake content regarding its source dataset or the", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1397 }, { "text": "corresponding facial manipulation techniques, and the malicious attacker is unlikely to reveal such crucial information. Consequently, for a victim of Deepfake to defend his or her innocence or accuse the attacker [22, 88, 97, 120], simply presenting a model detection decision and listing the numerical model performance on each benchmark dataset may not be convincing and reliable.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1398 }, { "text": "Specifically, a unique statistical claim is necessary regarding the detection performance to discuss the model trustworthiness on any arbitrary candidate suspect instead of varying on each testing dataset when adopting the detection model as forensic evidence for criminal investigation and court-case judgments. To conclude, the following questions are to be solved:", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1399 }, { "text": "\u2022 Can a detection model assist or act as evidence in forensic investigations in court? \u2022 How reliable is the detection model when performing as forensic evidence in real-life scenar- ios? \u2022 How accurate is the detection model regarding an arbitrary falsification? Unfortunately, to the best of our knowledge, no existing work has studied the model reliability", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1400 }, { "text": "or come up with a reliable claim for the model to play the role of forensic evidence. Therefore, in this study, inspired by the reliability study on antigen diagnostic tests [1, 27, 52] for the recent COVID-19 pandemic [132], we conduct a quantitative study by investigating the detection model reliability with a new evaluation metric with statistical techniques. Unlike the studies for antigen", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1401 }, { "text": "diagnostic tests that prefer to achieve a perfect specificity rather than sensitivity as the goal is to avoid missing any positive case [131], we wish to correctly identify both real and fake materials with no priority. In particular, we construct a population to imitate real-life Deepfake distribution", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1402 }, { "text": "and design a scientific random sampling scheme to analyze and compute the confidence intervals for the values of accuracy and AUC score metrics regarding the Deepfake detection models. As a result, numerical ranges indicating the reliable model performance at 90% and 95% confidence levels can be derived for both accuracy and AUC score.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1403 }, { "text": "4.2 Deepfake Population In reality, a candidate Deepfake suspect can only have two possible categories, namely, real and fake. Admittedly, most public images and videos circulating on the internet are pristine without artificial changes. However, whenever a real-life Deepfake case is raised such that the authenticity", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1404 }, { "text": "of the candidate material needs to be justified, we could not consider all images and videos in the world as the target population [166] because most of the real ones are not likely to be disputed in the discussion of Deepfake. At the same time, the probability that the candidate material is fake does not necessarily equal to the proportion of fake ones regarding all images and videos in the", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1405 }, { "text": "world. Despite the uncertainty of the real-life Deepfake population and distribution, in this study, we construct a sampling frame [12] with the accessible high-quality Deepfake benchmark datasets to imitate the target population of Deepfake in real life for detection model reliability analysis. Details", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1406 }, { "text": "12 Wang et al. 4.3 Random Sampling We perform random sampling [117] from the constructed sampling frame with a sample size of \ud835\udc60for \ud835\udc61trials. Two sampling options are considered in this study: balanced and imbalanced. For a balanced sampling setting, we maintain the condition that the same amount of real and fake", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1407 }, { "text": "samples are randomly drawn. On the contrary, an imbalanced setting allows a completely random stochastic rule with respect to real and fake samples. For an arbitrary Deepfake detection model \ud835\udc40after sufficient training, we randomly draw \ud835\udc60 samples from the sampling frame following the sampling option. Then the \ud835\udc60samples are fed to", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1408 }, { "text": "model \ud835\udc40for authenticity prediction, deriving predicted labels and prediction scores. After that, the accuracy and AUC score metrics are computed accordingly based on the ground-truth authenticities of the sampled faces. Such a sampling process is repeated for \ud835\udc61trials and a total of \ud835\udc61accuracies and \ud835\udc61AUC scores are derived in the end. Take the \ud835\udc61accuracies as an example, we first compute the", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1409 }, { "text": "mean value \u00af\ud835\udc65and standard deviation \ud835\udf0efollowing \u00af\ud835\udc65= \u03a3\ud835\udc61 \ud835\udc56=1\ud835\udc65\ud835\udc56 \ud835\udc61 , (1) and \ud835\udf0e= \u221a\ufe04 \u03a3\ud835\udc61 \ud835\udc56=1(\ud835\udc65\ud835\udc56\u2212\u00af\ud835\udc65)2 \ud835\udc61\u22121 , (2) where \ud835\udc65\ud835\udc56refers to the accuracy value of the \ud835\udc56-th trial. According to the central limit theorem (CLT) [45], the distribution of sample means tends toward a normal distribution as the sample size gets larger. Therefore, the normal distribution confidence", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1410 }, { "text": "interval \ud835\udc36\ud835\udc3ccan be calculated by \ud835\udc36\ud835\udc3c= \u00af\ud835\udc65\u00b1 \ud835\udc67\ud835\udf0e\u221a\ud835\udc60, (3) where parameter \ud835\udc67represents the z-score, an indicator of the confidence level following the instruction of the z-table [16]. When deducing statistical results for the \ud835\udc61AUC scores, the above workflow applies identically. Since the target population is of unknown distribution, different values of the sample size \ud835\udc60", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1411 }, { "text": "are adopted in order to settle the confidence intervals at different confidence levels. Meanwhile, considering that an insufficient number of trials per sample size may cause bias when locating the sample mean, different values of trails for \ud835\udc61are chosen to eliminate the potential bias. Detailed steps of the model reliability study can be summarized as Algorithm 1.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1412 }, { "text": "5 DATASET PREPARATION The choice of training dataset and data pre-processing scheme can significantly affect the per- formance of a deep learning model. Among the evolution of Deepfake datasets, as introduced in Section 2.2, various benchmark datasets are frequently adopted for training and testing to boost the", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1413 }, { "text": "detection model performance, but the workflow of data pre-processing operations has been barely discussed in detail in existing studies. Moreover, there lacks a standard pre-processing scheme in the current domain, causing difficulty in model comparison due to non-uniform training datasets after pre-processing by different detection work. On the other hand, using heedlessly prepared", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1414 }, { "text": "Deepfake Detection: A Comprehensive Survey from the Reliability Perspective 13 Algorithm 1 Deepfake detection model reliability study. Input: \ud835\udc40= well-trained detection model Output: 90% and 95% \ud835\udc36\ud835\udc3cfor each sample size 1: \ud835\udc53\ud835\udc5f\u2190list of real samples 2: \ud835\udc53\ud835\udc53\u2190list of fake samples 3: \ud835\udc5c\u2190balance sampling option", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1415 }, { "text": "4: \ud835\udc61\u2190number of trials 5: \ud835\udc46\u2190[\ud835\udc601,\ud835\udc602, ..., len(\ud835\udc46)] 6: shuffle(\u00b7) \u2190function to shuffle the list 7: acc(\u00b7) \u2190function to calculate accuracy 8: auc(\u00b7) \u2190function to calculate AUC score 9: for \ud835\udc56\u21900 to len(\ud835\udc46) \u22121 do 10: acc_lst \u2190[] 11: auc_lst \u2190[] 12: for \ud835\udc57\u21900 to \ud835\udc61do 13: if \ud835\udc5c== True then 14: \ud835\udc53\ud835\udc5f\u2190shuffle(\ud835\udc53\ud835\udc5f) 15: \ud835\udc53\ud835\udc53\u2190shuffle(\ud835\udc53\ud835\udc53)", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1416 }, { "text": "16: samples \u2190\ud835\udc53\ud835\udc5f[0 : \ud835\udc46[\ud835\udc56] 2 ] + \ud835\udc53\ud835\udc53[0 : \ud835\udc46[\ud835\udc56] 2 ] 17: else 18: \ud835\udc53\ud835\udc4e\u2190shuffle(\ud835\udc53\ud835\udc5f+ \ud835\udc53\ud835\udc53) 19: samples \u2190\ud835\udc53\ud835\udc4e[0 : \ud835\udc46[\ud835\udc56]] 20: end if 21: preds, pred_scores, labels \u2190\ud835\udc40(samples) 22: acc_lst.append(acc(preds, labels)) 23: auc_lst.append(auc(pred_scores, labels)) 24: end for 25: compute \u00af\ud835\udc65and \ud835\udf0efor acc_lst and auc_lst 26:", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1417 }, { "text": "compute and record 90% and 95% \ud835\udc36\ud835\udc3c 27: end for research domain, ensuring a fair game for other work to compare with the model performance as exhibited in this paper following the same settings. 5.1 Dataset Pre-processing While a video-level detector may rely on special data processing arrangement directly upon videos,", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1418 }, { "text": "for the frame-level detectors, since the detection results are evaluated based on all selected frames, to avoid potential biases toward particular videos, it is meaningful to keep the amount of extracted faces from each video equivalent during model training and testing. Therefore, we firstly obtain \ud835\udc50", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1419 }, { "text": "image frames using FFmpeg [44] for each candidate video with an equal frame interval between every two adjacent extracted frames following \ud835\udc43= \ud835\udc56\ud835\udc41 \ud835\udc50for 0 \u2264\ud835\udc56< \ud835\udc50and \ud835\udc56\u2208Z, (4) where \ud835\udc41refers to the number of frames that contain faces in the video and \ud835\udc43= {\ud835\udc5d0, \ud835\udc5d1, ..., \ud835\udc5d\ud835\udc50\u22121} contains the sequentially ordered indices for which frames to be extracted from the video. In other", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1420 }, { "text": "14 Wang et al. Besides, videos with fewer than \ud835\udc50frames containing detected faces are also omitted. The dlib library [90] is utilized for face detection and cropping where the face detector provides coordinates of the bounding box that locates the detected face. For the sequence of frames with frame indices", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1421 }, { "text": "\ud835\udc43= {\ud835\udc5d0, \ud835\udc5d1, ..., \ud835\udc5d\ud835\udc50\u22121} from a video, we fix the size \ud835\udc59of a squared bounding box \ud835\udc4ffor all faces by \ud835\udc59= max{max \ud835\udc56 \ud835\udc64\ud835\udc56, max \ud835\udc56 \u210e\ud835\udc56} for 0 \u2264\ud835\udc56< \ud835\udc50and \ud835\udc56\u2208Z, (5) where \ud835\udc64\ud835\udc56and \u210e\ud835\udc56are the widths and heights of each bounding box. We then locate the center of each face \ud835\udc53\ud835\udc56with the help of the corresponding bounding box \ud835\udc4f\ud835\udc56and place the fixed squared bounding", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1422 }, { "text": "box \ud835\udc4fat the centers for face cropping. 5.2 Datasets Involved and Detailed Arrangements Following the convention of the existing Deepfake detection work and considering the qual- ities of available benchmark datasets, we consider five datasets in experiments in this study, namely, FF++ [144], FaceShifter [100], DFDC [37], Celeb-DF [104], and DF1.0 [81]. In detail, early", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1423 }, { "text": "datasets [186] are excluded due to low quantity and diversity. Meanwhile, although WDF [197] is similar to real-life Deepfake materials, the videos collected from the internet are manually labeled without knowing the ground-truth labels, which leads to credibility issues. KoDF [95] is the largest Deepfake dataset up to date, but its huge magnitude requires unreasonably large", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1424 }, { "text": "storage (\u223c4 TB) that we are unable to acquire and process5. All involved benchmark datasets follow the pre-processing scheme for face extraction as discussed in Section 5.1, and special settings are mentioned in the following subsections when necessary. 5.2.1 FaceForensics++. FaceForensics++ (FF++) is currently the most widely adopted dataset in the", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1425 }, { "text": "existing Deepfake detection studies. The dataset contains 1,000 real videos collected from YouTube and 4,000 Deepfake videos synthesized based on the real ones. In specific, four facial manipulation techniques are each applied to the 1,000 real videos to derive the corresponding 1,000 fake ones. Among the four facial manipulation techniques, FaceSwap (FS) [115] and Deepfakes (DF) [34] are", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1426 }, { "text": "face-swapping algorithms that synthesize the faces by swapping facial identities, while Face2Face (F2F) [160] and NeuralTextures (NT) [159] perform face reenactment by modifying facial attributes such as expressions and accessories. The FF++ dataset has provided a subject-independent official dataset split with a ratio of", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1427 }, { "text": "720:140:140 for training, validation, and testing. Meanwhile, three dataset qualities have been released, namely, Raw, HQ (c23), and LQ (c40), where the latter two are compressed with different video compression levels following the H.264 codec. In recent Deepfake detection work, FF++ is frequently adopted as the training dataset due to its manipulation diversity and data orderliness,", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1428 }, { "text": "and the HQ (c23) version is mostly utilized because it has a similar video compression level and video quality to the real-life Deepfake contents. In this survey, whenever necessary, we adopt FF++ for model training following the official dataset split. The key image frames are also extracted and employed since the performance enhancement by the participating key image frames in the", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1429 }, { "text": "training process has been proved in the early studies [2, 104, 172]. In the training process, unless specially designed, commonly used data augmentation is performed upon the real faces to construct a balanced training dataset for real and fake. In the testing phase, the testing set is constructed following the official split without further augmentation.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1430 }, { "text": "5.2.2 Deepfake Detection Challenge. Deepfake Detection Challenge (DFDC) is one of the largest public Deepfake datasets with 23,654 real videos and 104,500 fake ones. Among the fake videos, 5The 6 manipulation algorithms in KoDF are highly overlapped with the 8 manipulation algorithms in DFDC. This favorably", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1431 }, { "text": "Deepfake Detection: A Comprehensive Survey from the Reliability Perspective 15 there are eight synthetic techniques [34, 73, 86, 129, 139, 187] that have been applied based on the real ones. Due to its large data quantity, we randomly pick 10 of the 50 video folders from the official dataset and randomly shuffle 100 real videos and 100 fake ones from each folder. Since", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1432 }, { "text": "most existing approaches focus on detecting Deepfake visually and most benchmark datasets are published without audio, fake videos using the pure audio swap technique are easily classified as pristine because there is no visual artifact on the faces. Meanwhile, the official DFDC dataset only provides labels for real and fake while sub-labels for specific synthetic techniques are unavailable.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1433 }, { "text": "Therefore, in this paper, the randomly picked 1,000 fake videos are manually examined to eliminate fake videos with the pure audio swap technique to omit noises in detection and guarantee a fair experimental setting. The selected videos are then fed through the data pre-processing scheme in Section 5.1 for model evaluation.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1434 }, { "text": "5.2.3 Celeb-DF. Celeb-DF is one of the most challenging benchmark Deepfake datasets that are publicly available. It contains 590 celebrity interview videos collected from YouTube and 5,639 face-swapped videos based on the real ones using an improved face-swapping algorithm with resolution enhancement, mismatched color correction [142], inaccurate face mask adjustment, and", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1435 }, { "text": "temporal flickering reducing [83] on the basic face-swapping auto-encoder architecture. A set of 518 official testing videos with high visual quality has failed most of the existing baseline models at a time because obvious visual artifacts can be barely found. We resort to the official testing set", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1436 }, { "text": "with 178 real videos and 340 fake ones for model evaluation. 5.2.4 DeeperForensics-1.0. DeeperForensics-1.0 (DF1.0) is the first large-scale dataset that is man- ually added with deliberate distortions and perturbations to the clean face-swapped videos. A strengthened face-swapping algorithm, Deepfake Variational Auto-Encoder (DF-VAE), is intro-", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1437 }, { "text": "duced for superior synthetic performance with better reenactments on expression and pose, fewer style mismatches, and more stable temporal continuity. There are a total of 10,000 synthesized videos where 1,000 of them are face-swapped from lab-controlled source videos onto the FF++ real videos using DF-VAE and the rest 9,000 videos are derived using the 1,000 raw manipulated", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1438 }, { "text": "videos by applying combinations of seven distortions6 under five intensity levels. Since the HQ and LQ versions of FF++ contain the same visual content and only differ in compression levels, DF1.0 with sufficient visual quality diversity in the manipulated videos serves as a perfect substitution for the LQ version of FF++ in the experiments, providing a convincing evaluation of all quality", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1439 }, { "text": "circumstances. Produced based on FF++, the dataset has only provided the official split ratio of 7:1:2 for the fake videos, and we thus execute model evaluation with merely the 2,000 fake testing videos. 5.2.5 FaceShifter. FaceShifter refers to a subject-agnostic GAN-based face-swapping algorithm that solves the facial occlusion challenge with a novel Heuristic Error Acknowledging Refinement", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1440 }, { "text": "Network (HEAR-Net). A subset with 1,000 synthetic videos is later included in the FF++ dataset by applying the FaceShifter face-swapping model to the 1,000 real videos. Since FaceShifter and FF++ share the same set of real videos, we take only the 140 fake videos for model evaluation following the FF++ official split.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1441 }, { "text": "6 DETECTION MODEL EVALUATION In this section, we first adopt several state-of-the-art Deepfake detection models that are mainly designed regarding each of the three challenges as defined in Section 3 and report their detection 6Change of color saturation, local block-wise distortion, change of color contrast, Gaussian blur, white Gaussian noise in", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1442 }, { "text": "16 Wang et al. performance on each benchmark testing set. Then the models are further discussed regarding the reliability following Algorithm 1 along with case studies on real-life Deepfake materials. 6.1 Experiment Settings Based on the Deepfake detection developing history and the three challenges of the current Deep-", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1443 }, { "text": "fake detection domain, in the experiment, we selected several representative milestone baseline models and the most recent ones that have source code publicly available for reproduction. Specifi- cally, Xception [24], MAT [192], and RECCE [13] mainly attempt on the transferability challenge, Stage5 [189] and FSTMatching [39] focus on the interpretability topic, and MetricLearning [14] and", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1444 }, { "text": "LRNet [156] are designed for the robustness issue. Models with publicly available trained weights are directly adopted for evaluation if the model is trained on FF++ or special arrangements other than the five benchmark datasets are necessary during training. The rest models are trained on FF++ in our experiment as discussed in Section 5 and all models converge commonly. The selected", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1445 }, { "text": "models are tested on all benchmark datasets. To guarantee complete fairness, we applied optimal parameter settings as reported in the corresponding published papers during training and testing. During model testing, we recorded the video-level Deepfake detection performance. In particular, detection results of the cropped faces of each video are averaged to a unique output for detectors", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1446 }, { "text": "that are designed for detecting individual images. Methods that can directly generate a single output for each video are fed with raw videos via the corresponding processing scheme as provided by their published source codes. The well-trained models are firstly evaluated on the FF++ testing set for the in-dataset setting, in other words, tested on the same dataset they have seen during training.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1447 }, { "text": "Then, to further validate the model performance on unseen datasets, the cross-dataset evaluation is conducted to test the models on DFDC, Celeb-DF, DF1.0, and FaceShifter. We set \ud835\udc41= 10 for training and \ud835\udc41= 20 for testing regarding Eq. (4) for frame extraction during data pre-processing when applicable.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1448 }, { "text": "We adopted accuracy (ACC) and AUC score at the video level as the evaluation metrics. In detail, the accuracy refers to the proportion of the correctly classified data items regarding all testing data, and the AUC score represents the area under the receiver operating characteristic (ROC) curve. In", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1449 }, { "text": "other words, the AUC score demonstrates the probability that a random positive sample scores higher than a random negative sample from the testing set, that is, the ability of the classifier to distinguish between real and fake faces. For testing sets that contain only fake samples, the AUC score is inapplicable and thus withdrawn.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1450 }, { "text": "6.2 Model Performance on Benchmark Testing Sets Model performance for both in- and cross-dataset evaluations is listed in Table 2. It can be observed that all models of the transferability topic have derived reasonable detection performance on FF++ with accuracy values and AUC scores over 90% since their goal is to achieve better performance in", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1451 }, { "text": "cross-dataset experiments after maintaining promising performance on seen data. In particular, MAT [192] wins the comparison with the highest 97.40% accuracy and 99.67% AUC score. On the contrary, models that are designed specifically for interpretability or robustness purposes have exhibited relatively poor detection performance on FF++, and the potential causation is discussed", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1452 }, { "text": "together with their detection performance on other benchmark datasets in the following paragraphs. As for cross-dataset evaluation, most models have suffered a performance damping since the testing data are unseen during training. In specific, no model has reached over 80% AUC scores on DFDC or Celeb-DF and some models even exhibit abnormal detection performance when", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1453 }, { "text": "validated on unseen fake testing sets solely (DF1.0 and FaceShifter). This may be caused by oblivious overfitting on real or fake data by the models. While models that are solving the transferability challenge have all exhibited normal and reasonable functionalities, hidden trouble can be discovered", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1454 }, { "text": "Deepfake Detection: A Comprehensive Survey from the Reliability Perspective 17 regarding the interpretability and robustness topics. Specifically, model weights of Stage5 [65] are adopted for testing because the model is trained with re-synthesized samples using exclusively GAN models under special settings, but this at the same time has led to unsatisfied results when", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1455 }, { "text": "detecting fake samples that are not synthesized using GAN architectures. FSTMatching [39] spends huge computing power on disentangling source and target artifacts for the explanation, which thus results in the failure against other models although similarly trained on FF++. MetricLearning [14] and LRNet [156] both are proposed and trained to deal with highly compressed Deepfake contents", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1456 }, { "text": "in special scenarios. Unfortunately, they suffer performance fluctuation when the compression condition varies without expectation. Furthermore, LRNet [156] executes detection based on facial landmarks solely, which is another main reason that leads to the unsatisfactory. Besides, models are generally unstable on different testing sets. For instance, RECCE [13] wins", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1457 }, { "text": "the competition on Celeb-DF for both accuracy and AUC score, but its detection ability deteriorates immensely when facing DFDC. On the other hand, MAT [192] wins the battle on DFDC and achieves competitive performance on Celeb-DF, but an obvious overdependence on the real samples can be concluded from its poor accuracies on DF1.0 and FaceShfiter. Meanwhile, Xception has derived", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1458 }, { "text": "reasonable detection performance on each testing set even though not winning the comparison on any dataset. As a result, no model appears to be the overall winner according to Table 2 and it is hard to determine which model to use when facing an arbitrary candidate Deepfake suspect in real-life cases.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1459 }, { "text": "Moreover, in most cases, a well-trained model usually achieves a higher AUC score than the accuracy on each testing set. The reason is that the threshold to classify real and fake with softmax or sigmoid function applied is always fixed at 0.5 for the accuracy evaluation upon the output scores within the range of [0, 1] where 0 refers to real and 1 represents fake, while the actual threshold for", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1460 }, { "text": "the optimal model performance is usually located differently regarding 0.5. Therefore, despite a classifier with a threshold value set to 0.5 does not perform well, the model may still distinguish between real and fake with a relatively high AUC score. However, it is also worth noting that although a high AUC score may reveal the model\u2019s ability to separate real and fake samples, the", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1461 }, { "text": "threshold may vary depending on different testing sets and different models. Hence, in order to stably determine real or fake, finding a fixed threshold to consistently satisfy the detection goal on arbitrary images and videos may help boost the overall model detection ability in the research domain.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1462 }, { "text": "65.21% 71.57% 70.27% 70.71% 56.87% 57.55% MAT [192] 97.40% 99.67% 66.63% 74.83% 71.81% 77.16% 41.74% 18.71% RECCE [13] 90.72% 95.26% 62.06% 66.94% 71.81% 77.90% 51.21% 56.12% Stage5\u2020 [65] 19.97% 50.21% 51.02% 48.08% 34.36% 39.88% 0.00% 0.00% FSTMatching [39] 81.33% 77.01% 44.88% 39.90% 38.61% 44.27%", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1463 }, { "text": "18 Wang et al. Table 3. Dataset statistics of the sampling frame for model testing regarding the number of videos with cropped faces. Datasets with no real samples are marked with \u2018\u2013\u2019 sign. FF++ [144] DFDC [37] Celeb-DF [104] DF1.0 [81] FaceShifter [100] Total Num Real 140 1,000 178 \u2013 \u2013 1,318 Num Fake", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1464 }, { "text": "560 1,000 340 2,010 140 4,050 Total 700 2,000 518 2,010 140 5,368 6.3 Model Reliability Experiment Results The model reliability evaluation is conducted on a sampling frame with 5,368 videos composed of the testing set of each benchmark dataset, and the detailed data quantity is listed in Table 3. Following the workflow of Algorithm 1, we proposed reliability analyses on the well-trained", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1465 }, { "text": "Deepfake detection approaches. We computed the 90% and 95% confidence intervals in experiments. Specifically, two values of \ud835\udc61, 500 and 3,000, are chosen to ensure a sufficient number of trials when locating the sample means. Various sample sizes \ud835\udc60\u2208{10, 100, 500, 1,000, 1,500, 2,000, 2,500} are selected to find the settled confidence intervals. Besides, both sampling options, balanced and", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1466 }, { "text": "imbalanced sampling, are executed in experiments. Detailed results are listed in Tables 4 to 7 following the Cartesian product settings of balancing options \ud835\udc42= {True, False}, the number of trials \ud835\udc47= {500, 3,000}, and evaluation metrics \ud835\udc38= {ACC, AUC}. It can be easily observed that for all models in the tables, the mean values \u00af\ud835\udc65gradually settle", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1467 }, { "text": "as sample sizes become larger and the standard deviation (Std.) values \ud835\udf0econsistently decrease concurrently. Similarly, the 90% and 95% confidence intervals are progressively straitened and settled around the mean values. Moreover, statistical results with 3,000 sampling trials converge faster and are more stable than those with 500 trials as sample size increases for all experiments,", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1468 }, { "text": "and both trial numbers lead to similar final values once settled. Regarding the balanced sampling setting, the results generally match the model detection perfor- mance in Table 2. On the other hand, since the constructed sampling frame is imbalanced for real and fake faces, an imbalanced sampling option may lead to more fake data than the real ones in the", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1469 }, { "text": "sample set. Consequently, for models that have shown relatively poor performance when evaluated on fake testing sets (DF1.0 and FaceShifter) solely, the mean values and confidence intervals of the accuracy are located at lower levels. Meanwhile, models that have achieved promising performance on merely the fake testing sets have led to higher accuracy values for mean and confidence intervals.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1470 }, { "text": "The AUC score metric, as mentioned in Section 6.1, is impervious regarding the imbalanced dataset. Therefore, mean values and confidence intervals under balanced and imbalanced sampling options are generally identical. With a closer look at the tables, the leading approaches, Xception [24] and MAT [192], both", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1471 }, { "text": "have achieved mean accuracies above 68% and confidence intervals around 69% with respect to the balanced sampling option. All models regarding the interpretability and robustness topics have derived accuracies and confidence intervals around 50%. Stage5 [65] and MetricLearning [14] convey results with all values being 50% since the former recognizes all candidate samples as real", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1472 }, { "text": "and the latter classifies all as fake when checking the predicted labels accordingly. While Stage5 [65] may only be interpretable when facing GAN-based synthetic contents depending on its model design, MetricLearning [14] relies on a fixed threshold of 5 upon the output value without softmax or sigmoid activation. Since a perfect threshold value may vary depending on the testing dataset,", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1473 }, { "text": "Deepfake Detection: A Comprehensive Survey from the Reliability Perspective 19 Table 4. Accuracy statistics (%) with balanced sampling for 500 and 3,000 trials. Model Names Num Trials Statistics Sample Sizes 10 100 500 1,000 1,500 2,000 2,500 Xception [24] 500 90% CI 57.89\u201379.55 65.67\u201372.21 67.40\u201370.24", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1474 }, { "text": "67.94\u201369.78 68.22\u201369.64 68.33\u201369.39 68.42\u201369.24 95% CI 55.81\u201381.63 65.04\u201372.84 67.13\u201370.51 67.76\u201369.96 68.08\u201369.78 68.23\u201369.49 68.34\u201369.31 Mean 68.72 68.94 68.82 68.86 68.93 68.86 68.83 Std. 14.70 4.44 1.92 1.25 0.97 0.72 0.56 3,000 90% CI 58.90\u201378.54 65.87\u201371.98 67.55\u201370.16 67.98\u201369.70 68.25\u201369.50", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1475 }, { "text": "68.37\u201369.36 68.50\u201369.26 95% CI 57.02\u201380.42 65.29\u201372.57 67.31\u201370.41 67.82\u201369.87 68.13\u201369.62 68.28\u201369.45 68.42\u201369.34 Mean 68.72 68.93 68.86 68.84 68.88 68.86 68.88 Std. 14.62 4.55 1.94 1.28 0.93 0.73 0.57 MAT [192] 500 90% CI 59.72\u201377.44 66.23\u201371.87 67.59\u201370.06 68.18\u201369.72 68.30\u201369.52 68.40\u201369.38 68.52\u201369.31", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1476 }, { "text": "95% CI 58.02\u201379.14 65.69\u201372.41 67.35\u201370.30 68.03\u201369.87 68.19\u201369.64 68.31\u201369.47 68.45\u201369.39 Mean 68.58 69.05 68.82 68.95 68.91 68.89 68.92 Std. 13.17 4.19 1.83 1.15 0.91 0.73 0.59 3,000 90% CI 59.45\u201377.74 66.10\u201371.84 67.67\u201370.15 68.04\u201369.71 68.28\u201369.50 68.38\u201369.37 68.49\u201369.29 95% CI 57.70\u201379.49 65.55\u201372.39", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1477 }, { "text": "67.43\u201370.39 67.88\u201369.87 68.16\u201369.62 68.29\u201369.47 68.41\u201369.37 Mean 68.60 68.97 68.91 68.87 68.89 68.88 68.89 Std. 13.61 4.27 1.85 1.24 0.91 0.74 0.60 RECCE [13] 500 90% CI 51.46\u201371.14 57.49\u201364.22 59.38\u201362.06 59.95\u201361.81 60.16\u201361.43 60.30\u201361.26 60.42\u201361.17 95% CI 49.56\u201373.04 56.84\u201364.86 59.13\u201362.32 59.77\u201361.99", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1478 }, { "text": "60.03\u201361.55 60.20\u201361.35 60.35\u201361.24 Mean 61.30 60.85 60.72 60.88 60.79 60.78 60.79 Std. 14.63 5.00 1.99 1.38 0.94 0.72 0.56 3,000 90% CI 49.56\u201372.51 57.19\u201364.21 59.28\u201362.27 59.78\u201361.76 60.08\u201361.52 60.25\u201361.35 60.37\u201361.22 95% CI 47.36\u201374.71 56.52\u201364.88 58.99\u201362.55 59.59\u201361.95 59.94\u201361.66 60.15\u201361.46", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1479 }, { "text": "60.29\u201361.30 Mean 61.04 60.70 60.77 60.77 60.80 60.80 60.80 Std. 15.59 4.77 2.03 1.35 0.98 0.75 0.58 Stage5 [65] 500 90% CI 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 95% CI 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 Mean", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1480 }, { "text": "50.00 50.00 50.00 50.00 50.00 50.00 50.00 Std. 0.00 0.00 0.00 0.00 0.00 0.00 0.00 3,000 90% CI 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 95% CI 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 Mean 50.00 50.00 50.00 50.00 50.00", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1481 }, { "text": "50.00 50.00 Std. 0.00 0.00 0.00 0.00 0.00 0.00 0.00 FSTMatching [39] 500 90% CI 44.61\u201359.71 49.85\u201354.76 51.42\u201353.37 51.68\u201353.01 51.91\u201352.87 51.96\u201352.78 52.06\u201352.69 95% CI 43.16\u201361.16 49.37\u201355.23 51.24\u201353.56 51.55\u201353.14 51.82\u201352.96 51.88\u201352.86 52.00\u201352.75 Mean 52.16 52.30 52.40 52.34 52.39 52.37 52.37", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1482 }, { "text": "Std. 12.12 3.95 1.56 1.07 0.77 0.66 0.50 3,000 90% CI 44.62\u201359.82 50.15\u201354.80 51.42\u201353.42 51.73\u201353.07 51.85\u201352.84 51.98\u201352.78 52.06\u201352.70 95% CI 43.16\u201361.28 49.71\u201355.24 51.23\u201353.61 51.60\u201353.19 51.76\u201352.94 51.91\u201352.85 51.99\u201352.76 Mean 52.22 52.48 52.42 52.40 52.35 52.38 52.38 Std. 12.23 3.73 1.60 1.07", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1483 }, { "text": "0.80 0.64 0.52 MetricLearning [14] 500 90% CI 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 95% CI 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 Mean 50.00 50.00 50.00 50.00 50.00 50.00 50.00 Std. 0.00 0.00 0.00 0.00 0.00 0.00", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1484 }, { "text": "0.00 3,000 90% CI 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 95% CI 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 50.00\u201350.00 Mean 50.00 50.00 50.00 50.00 50.00 50.00 50.00 Std. 0.00 0.00 0.00 0.00 0.00 0.00 0.00 LRNet [156] 500 90% CI", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1485 }, { "text": "41.70\u201362.46 48.22\u201355.13 50.53\u201353.24 51.09\u201352.97 51.31\u201352.72 51.42\u201352.51 51.62\u201352.37 95% CI 39.70\u201364.46 47.56\u201355.79 50.27\u201353.51 50.91\u201353.15 51.18\u201352.86 51.32\u201352.61 51.55\u201352.44 Mean 52.08 51.67 51.89 52.03 52.02 51.96 52.00 Std. 15.43 5.13 2.02 1.40 1.05 0.81 0.55 3,000 90% CI 41.70\u201361.81 49.05\u201355.25", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1486 }, { "text": "20 Wang et al. Table 5. Accuracy statistics (%) with imbalanced sampling for 500 and 3,000 trials. Model Names Num Trials Statistics Sample Sizes 10 100 500 1,000 1,500 2,000 2,500 Xception [24] 500 90% CI 55.90\u201375.50 62.70\u201369.19 64.66\u201367.62 65.08\u201367.19 65.48\u201367.06 65.55\u201366.91 65.58\u201366.89 95% CI 54.02\u201377.38", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1487 }, { "text": "62.07\u201369.82 64.37\u201367.90 64.88\u201367.39 65.32\u201367.21 65.42\u201367.05 65.45\u201367.02 Mean 65.70 65.94 66.14 66.13 66.27 66.23 66.23 Std. 14.56 4.83 2.20 1.57 1.18 1.01 0.97 3,000 90% CI 55.36\u201377.38 62.75\u201369.79 64.74\u201367.81 65.06\u201367.25 65.32\u201367.06 65.44\u201366.99 65.51\u201366.94 95% CI 53.24\u201379.50 62.08\u201370.46 64.44\u201368.10", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1488 }, { "text": "64.85\u201367.46 65.15\u201367.23 65.29\u201367.13 65.37\u201367.07 Mean 66.37 66.27 66.27 66.15 66.19 66.21 66.22 Std. 14.97 4.78 2.09 1.49 1.19 1.05 0.97 MAT [192] 500 90% CI 50.11\u201370.45 57.75\u201364.32 59.17\u201361.98 59.61\u201361.72 59.82\u201361.55 59.93\u201361.38 59.93\u201361.28 95% CI 48.15\u201372.41 57.12\u201364.95 58.90\u201362.24 59.41\u201361.92 59.65\u201361.72", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1489 }, { "text": "59.79\u201361.52 59.80\u201361.41 Mean 60.28 61.03 60.57 60.66 60.68 60.66 60.61 Std. 15.12 4.88 2.08 1.57 1.29 1.08 1.00 3,000 90% CI 49.79\u201370.58 57.22\u201363.89 59.21\u201362.13 59.57\u201361.64 59.80\u201361.52 59.90\u201361.39 59.97\u201361.26 95% CI 47.80\u201372.57 56.59\u201364.53 58.94\u201362.40 59.37\u201361.83 59.64\u201361.69 59.76\u201361.53 59.85\u201361.38", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1490 }, { "text": "Mean 60.18 60.56 60.67 60.60 60.66 60.64 60.61 Std. 15.47 4.96 2.17 1.54 1.28 1.11 0.96 RECCE [13] 500 90% CI 53.35\u201373.89 59.48\u201366.01 61.53\u201364.47 62.15\u201364.17 62.37\u201364.00 62.30\u201363.70 62.43\u201363.76 95% CI 51.37\u201375.87 58.86\u201366.64 61.25\u201364.75 61.96\u201364.36 62.21\u201364.16 62.16\u201363.84 62.30\u201363.88 Mean 63.62 62.75", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1491 }, { "text": "63.00 63.16 63.18 63.00 63.09 Std. 15.27 4.85 2.18 1.50 1.21 1.04 0.98 3,000 90% CI 52.85\u201373.05 59.88\u201366.32 61.64\u201364.53 62.01\u201364.07 62.25\u201363.89 62.30\u201363.77 62.36\u201363.66 95% CI 50.91\u201374.99 59.26\u201366.94 61.36\u201364.81 61.81\u201364.27 62.09\u201364.04 62.16\u201363.91 62.24\u201363.79 Mean 62.95 63.10 63.09 63.04 63.07 63.04", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1492 }, { "text": "63.01 Std. 15.04 4.80 2.15 1.54 1.22 1.09 0.97 Stage5 [65] 500 90% CI 16.38\u201333.50 22.20\u201327.30 23.73\u201326.06 24.04\u201325.65 24.20\u201325.34 24.33\u201325.28 24.37\u201325.15 95% CI 14.73\u201335.15 21.71\u201327.79 23.50\u201326.28 23.88\u201325.80 24.10\u201325.45 24.24\u201325.37 24.30\u201325.23 Mean 24.94 24.75 24.89 24.84 24.77 24.80 24.76 Std. 13.75", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1493 }, { "text": "4.09 1.87 1.29 0.91 0.76 0.63 3,000 90% CI 16.38\u201333.20 22.13\u201327.53 23.61\u201325.96 24.02\u201325.54 24.21\u201325.37 24.31\u201325.25 24.40\u201325.17 95% CI 14.77\u201334.82 21.62\u201328.05 23.39\u201326.18 23.88\u201325.69 24.10\u201325.49 24.22\u201325.34 24.33\u201325.25 Mean 24.79 24.83 24.79 24.78 24.79 24.78 24.79 Std. 13.52 4.34 1.89 1.22 0.93 0.76", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1494 }, { "text": "0.62 FSTMatching [39] 500 90% CI 31.38\u201351.22 37.87\u201344.02 39.65\u201342.21 40.20\u201341.92 40.43\u201341.76 40.56\u201341.64 40.67\u201341.54 95% CI 29.47\u201353.13 37.28\u201344.61 39.40\u201342.45 40.03\u201342.08 40.31\u201341.88 40.45\u201341.74 40.58\u201341.62 Mean 41.30 40.95 40.93 41.06 41.09 41.10 41.10 Std. 15.93 4.94 2.06 1.38 1.06 0.87 0.70 3,000", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1495 }, { "text": "90% CI 31.47\u201350.66 38.02\u201344.20 39.83\u201342.47 40.19\u201341.94 40.42\u201341.74 40.56\u201341.64 40.65\u201341.53 95% CI 29.63\u201352.50 37.43\u201344.79 39.58\u201342.72 40.02\u201342.11 40.29\u201341.87 40.45\u201341.75 40.57\u201341.62 Mean 41.07 41.11 41.15 41.06 41.08 41.10 41.09 Std. 15.43 4.97 2.12 1.41 1.06 0.87 0.71 MetricLearning [14] 500 90% CI", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1496 }, { "text": "66.17\u201383.91 72.88\u201377.98 74.09\u201376.40 74.42\u201376.06 74.65\u201375.78 74.69\u201375.67 74.86\u201375.59 95% CI 64.47\u201385.61 72.39\u201378.47 73.87\u201376.62 74.27\u201376.22 74.54\u201375.89 74.60\u201375.77 74.79\u201375.66 Mean 75.04 75.43 75.25 75.24 75.22 75.18 75.22 Std. 14.23 4.09 1.85 1.31 0.91 0.79 0.59 3,000 90% CI 66.53\u201383.60 72.81\u201378.12", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1497 }, { "text": "74.05\u201376.36 74.46\u201375.97 74.66\u201375.83 74.75\u201375.67 74.85\u201375.62 95% CI 64.89\u201385.24 72.30\u201378.63 73.83\u201376.58 74.32\u201376.12 74.54\u201375.95 74.66\u201375.76 74.77\u201375.69 Mean 75.07 75.46 75.20 75.22 75.25 75.21 75.23 Std. 13.73 4.28 1.86 1.22 0.95 0.74 0.62 LRNet [156] 500 90% CI 44.08\u201362.36 49.99\u201355.83 51.62\u201354.32 51.89\u201353.62", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1498 }, { "text": "52.08\u201353.41 52.14\u201353.19 52.19\u201353.07 95% CI 42.32\u201364.12 49.43\u201356.40 51.36\u201354.58 51.72\u201353.79 51.96\u201353.53 52.04\u201353.29 52.10\u201353.16 Mean 53.22 52.91 52.97 52.75 52.74 52.66 52.63 Std. 14.68 4.69 2.17 1.39 1.06 0.84 0.71 3,000 90% CI 42.91\u201362.33 49.72\u201355.84 51.33\u201353.99 51.79\u201353.58 51.99\u201353.36 52.13\u201353.22", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1499 }, { "text": "Deepfake Detection: A Comprehensive Survey from the Reliability Perspective 21 Table 6. AUC score statistics (%) with balanced sampling for 500 and 3,000 trials. Model Names Num Trials Statistics Sample Sizes 10 100 500 1,000 1,500 2,000 2,500 Xception [24] 500 90% CI 68.57\u201384.16 74.21\u201378.76 75.67\u201377.53", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1500 }, { "text": "76.06\u201377.32 76.18\u201377.13 76.29\u201377.04 76.38\u201376.92 95% CI 67.08\u201385.66 73.77\u201379.20 75.49\u201377.71 75.94\u201377.44 76.09\u201377.22 76.22\u201377.11 76.33\u201376.98 Mean 76.37 76.49 76.60 76.69 76.65 76.66 76.65 Std. 15.69 4.58 1.87 1.27 0.95 0.76 0.55 3,000 90% CI 69.14\u201384.68 74.50\u201379.02 75.69\u201377.64 76.02\u201377.29 76.23\u201377.14", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1501 }, { "text": "76.30\u201377.02 76.36\u201376.93 95% CI 67.65\u201386.17 74.07\u201379.45 75.51\u201377.82 75.90\u201377.42 76.14\u201377.23 76.23\u201377.09 76.30\u201376.98 Mean 76.91 76.76 76.66 76.66 76.69 76.66 76.64 Std. 15.67 4.55 1.96 1.28 0.92 0.73 0.57 MAT [192] 500 90% CI 66.83\u201383.07 73.00\u201377.79 74.79\u201376.75 74.88\u201376.21 75.13\u201376.16 75.31\u201376.12 75.31\u201375.93", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1502 }, { "text": "95% CI 65.27\u201384.63 72.54\u201378.25 74.60\u201376.94 74.75\u201376.34 75.03\u201376.26 75.23\u201376.20 75.25\u201375.99 Mean 74.95 75.39 75.77 75.54 75.64 75.71 75.62 Std. 16.34 4.81 1.97 1.34 1.05 0.82 0.62 3,000 90% CI 67.60\u201383.56 73.37\u201378.00 74.68\u201376.61 74.94\u201376.27 75.13\u201376.12 75.24\u201376.03 75.33\u201375.96 95% CI 66.07\u201385.09 72.93\u201378.45", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1503 }, { "text": "74.50\u201376.79 74.82\u201376.40 75.03\u201376.21 75.17\u201376.10 75.27\u201376.02 Mean 75.58 75.69 75.65 75.61 75.62 75.63 75.64 Std. 16.09 4.67 1.94 1.34 1.00 0.79 0.64 RECCE [13] 500 90% CI 57.27\u201375.43 63.82\u201369.22 65.88\u201368.13 66.02\u201367.53 66.18\u201367.32 66.39\u201367.27 66.40\u201367.14 95% CI 55.52\u201377.18 63.30\u201369.74 65.66\u201368.35 65.87\u201367.67", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1504 }, { "text": "66.07\u201367.43 66.31\u201367.35 66.33\u201367.21 Mean 66.35 66.52 67.00 66.77 66.75 66.83 66.77 Std. 18.28 5.43 2.26 1.52 1.14 0.88 0.75 3,000 90% CI 57.80\u201375.48 64.02\u201369.42 65.65\u201367.90 66.05\u201367.59 66.25\u201367.41 66.35\u201367.26 66.44\u201367.19 95% CI 56.11\u201377.17 63.50\u201369.94 65.43\u201368.12 65.90\u201367.73 66.14\u201367.52 66.26\u201367.35", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1505 }, { "text": "66.37\u201367.26 Mean 66.64 66.72 66.77 66.82 66.83 66.80 66.82 Std. 17.81 5.45 2.27 1.55 1.17 0.92 0.75 Stage5 [65] 500 90% CI 42.69\u201357.45 48.82\u201352.75 49.70\u201351.45 49.80\u201350.95 50.07\u201350.95 50.17\u201350.88 50.20\u201350.76 95% CI 41.27\u201358.86 48.44\u201353.12 49.53\u201351.62 49.68\u201351.07 49.99\u201351.03 50.10\u201350.95 50.15\u201350.81 Mean", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1506 }, { "text": "50.07 50.78 50.57 50.38 50.51 50.53 50.48 Std. 20.02 5.33 2.37 1.57 1.19 0.97 0.76 3,000 90% CI 43.18\u201357.79 48.23\u201352.59 49.59\u201351.37 49.81\u201351.03 50.06\u201350.96 50.13\u201350.83 50.22\u201350.75 95% CI 41.78\u201359.20 47.81\u201353.01 49.42\u201351.55 49.70\u201351.14 49.97\u201351.04 50.06\u201350.90 50.16\u201350.80 Mean 50.49 50.41 50.48 50.42", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1507 }, { "text": "50.51 50.48 50.48 Std. 19.36 5.78 2.36 1.61 1.19 0.94 0.71 FSTMatching [39] 500 90% CI 47.67\u201366.44 54.04\u201359.66 55.80\u201358.10 56.09\u201357.72 56.39\u201357.53 56.45\u201357.40 56.53\u201357.28 95% CI 45.87\u201368.24 53.49\u201360.20 55.58\u201358.32 55.93\u201357.88 56.29\u201357.64 56.36\u201357.49 56.46\u201357.35 Mean 57.06 56.85 56.95 56.91 56.96 56.92", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1508 }, { "text": "56.90 Std. 18.88 5.66 2.31 1.64 1.15 0.96 0.76 3,000 90% CI 47.06\u201366.20 54.10\u201359.72 55.84\u201358.24 56.16\u201357.78 56.27\u201357.49 56.45\u201357.37 56.56\u201357.31 95% CI 45.23\u201368.04 53.56\u201360.26 55.61\u201358.47 56.00\u201357.93 56.16\u201357.61 56.36\u201357.45 56.49\u201357.38 Mean 56.63 56.91 57.04 56.97 56.88 56.91 56.93 Std. 19.29 5.66 2.42", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1509 }, { "text": "1.63 1.23 0.92 0.75 MetricLearning [14] 500 90% CI 64.48\u201376.77 67.55\u201371.48 68.57\u201370.18 69.01\u201370.08 68.99\u201369.81 69.22\u201369.78 69.28\u201369.69 95% CI 63.29\u201377.95 67.18\u201371.85 68.41\u201370.33 68.90\u201370.19 68.91\u201369.89 69.17\u201369.83 69.24\u201369.73 Mean 70.62 69.51 69.37 69.55 69.40 69.50 69.49 Std. 16.26 5.19 2.13 1.42 1.09", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1510 }, { "text": "0.73 0.54 3,000 90% CI 60.11\u201378.39 66.82\u201372.23 68.35\u201370.58 68.78\u201370.23 68.91\u201369.96 69.10\u201369.90 69.20\u201369.78 95% CI 58.36\u201380.14 66.30\u201372.74 68.14\u201370.79 68.64\u201370.37 68.80\u201370.07 69.03\u201369.98 69.14\u201369.83 Mean 69.25 69.52 69.47 69.50 69.44 69.50 69.49 Std. 17.56 5.20 2.14 1.40 1.02 0.76 0.56 LRNet [156] 500", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1511 }, { "text": "90% CI 46.90\u201364.71 51.93\u201357.71 53.66\u201356.02 54.26\u201355.81 54.44\u201355.61 54.55\u201355.43 54.69\u201355.32 95% CI 45.19\u201366.43 51.37\u201358.27 53.43\u201356.25 54.11\u201355.96 54.32\u201355.73 54.46\u201355.52 54.63\u201355.38 Mean 55.81 54.82 54.84 55.04 55.03 54.99 55.01 Std. 17.93 5.83 2.38 1.57 1.19 0.89 0.64 3,000 90% CI 45.24\u201363.92 52.49\u201358.15", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1512 }, { "text": "22 Wang et al. Table 7. AUC score statistics (%) with imbalanced sampling for 500 and 3,000 trials. Model Names Num Trials Statistics Sample Sizes 10 100 500 1,000 1,500 2,000 2,500 Xception [24] 500 90% CI \u2013 74.64\u201378.75 75.80\u201377.60 75.98\u201377.21 76.12\u201377.19 76.17\u201377.06 76.31\u201377.15 95% CI \u2013 74.25\u201379.15", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1513 }, { "text": "75.62\u201377.77 75.86\u201377.33 76.02\u201377.29 76.09\u201377.15 76.23\u201377.23 Mean \u2013 76.70 76.70 76.60 76.65 76.62 76.73 Std. \u2013 5.43 2.38 1.62 1.41 1.18 1.11 3,000 90% CI \u2013 74.78\u201378.80 75.79\u201377.56 76.02\u201377.29 76.13\u201377.17 76.20\u201377.11 76.24\u201377.05 95% CI \u2013 74.39\u201379.19 75.61\u201377.73 75.89\u201377.41 76.03\u201377.27 76.11\u201377.20 76.17\u201377.13", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1514 }, { "text": "Mean \u2013 76.79 76.67 76.65 76.65 76.66 76.65 Std. \u2013 5.33 2.35 1.69 1.38 1.21 1.07 MAT [192] 500 90% CI \u2013 74.03\u201378.10 74.64\u201376.44 75.09\u201376.35 75.09\u201376.11 75.24\u201376.14 75.23\u201376.02 95% CI \u2013 73.64\u201378.49 74.47\u201376.62 74.97\u201376.47 74.99\u201376.21 75.15\u201376.23 75.16\u201376.09 Mean \u2013 76.07 75.54 75.72 75.60 75.69 75.63 Std.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1515 }, { "text": "\u2013 5.24 2.32 1.62 1.32 1.16 1.01 3,000 90% CI \u2013 73.65\u201377.52 74.71\u201376.41 75.07\u201376.28 75.16\u201376.16 75.25\u201376.10 75.25\u201376.03 95% CI \u2013 73.28\u201377.89 74.55\u201376.58 74.96\u201376.40 75.07\u201376.26 75.17\u201376.18 75.18\u201376.10 Mean \u2013 75.58 75.56 75.68 75.66 75.68 75.64 Std. \u2013 5.13 2.26 1.60 1.32 1.13 1.03 RECCE [13] 500 90% CI", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1516 }, { "text": "\u2013 64.69\u201369.07 65.67\u201367.67 66.32\u201367.58 66.27\u201367.36 66.35\u201367.31 66.37\u201367.26 95% CI \u2013 64.27\u201369.50 65.48\u201367.86 66.20\u201367.70 66.17\u201367.46 66.26\u201367.40 66.29\u201367.34 Mean \u2013 66.88 66.67 66.95 66.81 66.83 66.82 Std. \u2013 5.80 2.64 1.66 1.44 1.27 1.17 3,000 90% CI \u2013 64.67\u201368.97 65.85\u201367.78 66.17\u201367.56 66.26\u201367.38 66.34\u201367.28", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1517 }, { "text": "66.38\u201367.25 95% CI \u2013 64.26\u201369.38 65.66\u201367.96 66.04\u201367.69 66.16\u201367.48 66.25\u201367.37 66.30\u201367.33 Mean \u2013 66.82 66.81 66.87 66.82 66.81 66.81 Std. \u2013 5.69 2.56 1.84 1.47 1.25 1.15 Stage5 [65] 500 90% CI \u2013 48.58\u201353.40 49.39\u201351.49 49.74\u201351.14 50.00\u201351.09 50.08\u201350.98 50.13\u201350.83 95% CI \u2013 48.12\u201353.86 49.19\u201351.69", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1518 }, { "text": "49.61\u201351.27 49.89\u201351.19 50.00\u201351.07 50.07\u201350.90 Mean \u2013 50.99 50.44 50.44 50.54 50.53 50.48 Std. \u2013 6.37 2.78 1.85 1.44 1.18 0.92 3,000 90% CI \u2013 48.10\u201352.96 49.49\u201351.58 49.78\u201351.18 49.97\u201351.02 50.06\u201350.91 50.11\u201350.81 95% CI \u2013 47.64\u201353.42 49.29\u201351.78 49.65\u201351.31 49.87\u201351.12 49.98\u201350.99 50.04\u201350.88 Mean", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1519 }, { "text": "\u2013 50.53 50.53 50.48 50.49 50.49 50.46 Std. \u2013 6.44 2.77 1.85 1.39 1.13 0.93 FSTMatching [39] 500 90% CI \u2013 54.79\u201359.43 55.72\u201357.70 56.33\u201357.62 56.45\u201357.46 56.53\u201357.33 56.64\u201357.30 95% CI \u2013 54.34\u201359.87 55.53\u201357.89 56.20\u201357.74 56.35\u201357.56 56.45\u201357.40 56.57\u201357.37 Mean \u2013 57.11 56.71 56.97 56.96 56.93 56.97", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1520 }, { "text": "Std. \u2013 6.14 2.62 1.71 1.34 1.06 0.88 3,000 90% CI \u2013 54.53\u201359.18 56.08\u201358.08 56.21\u201357.55 56.42\u201357.44 56.52\u201357.36 56.60\u201357.27 95% CI \u2013 54.08\u201359.63 55.89\u201358.27 56.09\u201357.68 56.33\u201357.53 56.44\u201357.44 56.54\u201357.34 Mean \u2013 56.86 57.08 56.88 56.93 56.94 56.94 Std. \u2013 6.16 2.64 1.77 1.34 1.11 0.89 MetricLearning [14]", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1521 }, { "text": "500 90% CI \u2013 66.94\u201372.04 68.33\u201370.44 68.75\u201370.16 69.06\u201370.11 69.06\u201369.91 69.02\u201369.76 95% CI \u2013 66.45\u201372.53 68.13\u201370.64 68.62\u201370.29 68.96\u201370.21 68.98\u201369.99 68.95\u201369.83 Mean \u2013 69.49 69.38 69.46 69.59 69.48 69.39 Std. \u2013 6.74 2.78 1.85 1.38 1.12 0.97 3,000 90% CI \u2013 67.07\u201372.09 68.45\u201370.49 68.72\u201370.12 68.93\u201370.01", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1522 }, { "text": "69.07\u201369.91 69.16\u201369.87 95% CI \u2013 66.59\u201372.58 68.25\u201370.69 68.58\u201370.25 68.83\u201370.12 68.99\u201369.99 69.09\u201369.94 Mean \u2013 69.58 69.47 69.42 69.47 69.49 69.51 Std. \u2013 6.65 2.71 1.86 1.43 1.12 0.94 LRNet [156] 500 90% CI \u2013 52.42\u201357.52 54.03\u201356.28 54.33\u201355.77 54.56\u201355.71 54.48\u201355.37 54.59\u201355.42 95% CI \u2013 51.93\u201358.01", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1523 }, { "text": "53.81\u201356.50 54.19\u201355.91 54.45\u201355.81 54.40\u201355.46 54.51\u201355.50 Mean \u2013 54.97 55.15 55.05 55.13 54.93 55.01 Std. \u2013 6.56 2.90 1.85 1.47 1.14 1.06 3,000 90% CI \u2013 52.53\u201357.66 53.86\u201356.03 54.32\u201355.78 54.43\u201355.56 54.56\u201355.46 54.64\u201355.39 95% CI \u2013 52.04\u201358.16 53.65\u201356.24 54.17\u201355.92 54.32\u201355.67 54.47\u201355.54 54.57\u201355.46", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1524 }, { "text": "Deepfake Detection: A Comprehensive Survey from the Reliability Perspective 23 Fig. 2. Screenshots of the four real-life Deepfake videos for Deepfake detection in the case study. The videos are hyper-realistic with different resolutions and no obvious artifacts can be observed visually by human eyes.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1525 }, { "text": "real and fake with a mean value of 69.49% AUC score. As for FSTMatching [39] and LRNet [156], the experimental results are generally consistent with the ones in Table 2. As for experiments under the imbalanced sampling setting, besides the foreseeably high and low performance by MetricLearning [14] and Stage5 [65] as discussed, Xception [24] wins with", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1526 }, { "text": "the highest mean accuracy of 66.23% due to its stable performance in detecting both real and fake faces. Regarding the AUC scores, ignoring the imperceptible differences and taking a look at Table 6, 76.64% is derived by Xception because of its ability to separate real and fake samples at a certain threshold even though performing relatively unsatisfactory with the threshold value of 0.5", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1527 }, { "text": "regarding the softmax output scores. Besides, MAT [192] is the only other model that has achieved above 70% AUC score. In the remaining approaches, RECCE has reached above 65% AUC scores, while the other two methods have performed relatively unsatisfactory in comparison. It is also worth noting that despite the imbalanced sampling setting does not affect the final", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1528 }, { "text": "values of mean and confidence intervals, it may cause the AUC score incomputable for a tiny sample size. In particular, as shown in Table 7, there is a high possibility to randomly draw 10 samples that belong to the same category, which then leads to an incomputable AUC score since the sample set lacks data from the other category. Meanwhile, it is unlikely to randomly draw 100", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1529 }, { "text": "or even more samples that are of the same category even though there is a possibility theoretically. 7 REAL-LIFE CASE STUDY In this section, we made use of the experiment results of the model reliability study. Experiments are conducted to analyze the reliabilities of the existing Deepfake detection approaches when", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1530 }, { "text": "applied to real-life cases. Specifically, four famous Deepfake cases that have jeopardized individuals and society from 2018 to 2022 are considered in this case study (Fig. 2). In 2018, when the technique of Deepfake had just been released shortly, the well-known actress Emma Watson who performed in the Harry Potter movie series was face-swapped onto porn videos [88]. In the same year, another", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1531 }, { "text": "famous actress Natalie Portman encountered a similar fake scandal because of Deepfake [97]. The porn videos were widely spread at the time and had gravely influenced their reputations because the term \u2018Deepfake\u2019 was unfamiliar to the public and people were easily tricked and believed the videos to be genuine upon their first appearances. Later in March 2021, a Bucks County mom", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1532 }, { "text": "was accused of creating Deepfake videos of the underage girls on her daughter\u2019s cheerleader team and threatening them to quit the team [22]. The videos exhibited the girls that were naked, drinking alcohol, or vaping, and are accused to be fake. Nevertheless, two months later in May, the prosecutors admitted that they could not prove the fake-video claims without reliable tools", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1533 }, { "text": "24 Wang et al. Table 8. Deepfake detection results on real-life cases by the well-trained models. 95% confidence intervals following the balanced sampling option for accuracies are listed along the models. (\u2020: threshold fixed as 5 where greater values refers to fake; \u2021: video-level detector with no intermediate frame-level result.)", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1534 }, { "text": "Model Real-life Videos 95% CI Balanced Sampling ACC (%) Watson (2018) Portman (2018) Cheerleader (2021) Zelensky (2022) # Real / Fake Fake Score # Real / Fake Fake Score # Real / Fake Fake Score # Real / Fake Fake Score Xception [24] 78 / 22 0.290 (Real) 59 / 41 0.426 (Real) 68 / 7 0.401 (Real) 69 / 31", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1535 }, { "text": "0.328 (Real) 68.42\u201369.34 MAT [192] 66 / 34 0.365 (Real) 67 / 33 0.378 (Real) 3 / 72 0.951 (Fake) 34 / 66 0.630 (Fake) 68.41\u201369.37 RECCE [13] 10 / 90 0.779 (Fake) 8 / 92 0.808 (Fake) 0 / 75 0.902 (Fake) 15 / 85 0.723 (Fake) 60.29\u201361.30 Stage5 [65] 100 / 0 0.000 (Real) 100 / 0 0.000 (Real) 75 / 0 0.000 (Real)", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1536 }, { "text": "100 / 0 0.000 (Real) 50.00\u201350.00 FSTMatching [39] 94 / 6 0.259 (Real) 89 / 11 0.261 (Real) 75 / 0 0.113 (Real) 95 / 5 0.245 (Real) 51.99\u201352.76 MetricLearning\u2020 [14] 35 / 65 11.698 (Fake) 27 / 73 11.996 (Fake) 23 / 52 11.928 (Fake) 27 / 73 11.951 (Fake) 50.00\u201350.00 LRNet\u2021 [156] \u2013 0.655 (Fake) \u2013 0.667 (Fake)", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1537 }, { "text": "\u2013 0.667 (Fake) \u2013 0.538 (Fake) 51.55\u201352.43 video, a synthetic president Zelensky was telling Ukrainians to put down their weapons and give up resistance. 7.1 Detection Results We obtained the available video clips of the four Deepfake cases from the internet and performed Deepfake detection using each of the well-trained models. Specifically, for video clips with sufficient", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1538 }, { "text": "numbers of image frames that contain faces, we randomly extracted 100 frames for face cropping when using frame-level detectors. As for the cheerleader case, as the sensitive contents are omitted, we were only able to acquire a total of 75 faces from the news clip. For frame-level detectors, the numbers of faces classified as real or fake by each model for each video are exhibited in Table 8,", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1539 }, { "text": "and except MetricLearning [14] that uses a fixed threshold, the softmax scores for each video are averaged to obtain a single score that indicates the model determined authenticity. In particular, except for MetricLearning, the fake scores lie in the range of [0, 1] where 0 refers to real and 1 refers to fake. For video-level detectors, the ultimate results are straightforwardly demonstrated", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1540 }, { "text": "in the table. As a result, given the fact that all four videos are known to be fake, about half of the selected models have made the correct classifications regarding both the number of fake faces and the average softmax score. Besides, we provided reliably quantified 95% confidence intervals regarding the models\u2019 detection accuracy for reference when utilizing the results.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1541 }, { "text": "Despite achieving high statistical values regarding the accuracy and AUC score metrics in early experiments, Xception has failed to classify all four fake videos such that most faces are classified as real and all average softmax scores are below 0.43. Besides, the fake Emma Watson and Natalie Portman videos have also tricked the MAT model such that roughly two-thirds of the faces are", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1542 }, { "text": "classified as real and the final scores are below 0.4. Stage5 and FSTMatching, which proved to be relatively unsatisfactory in early experiments and discussions, have both failed to detect all four fake videos. MetricLearning and LRNet, although performing poorly on lab-controlled datasets, have shown robust detection ability especially because the videos circulated and downloaded from", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1543 }, { "text": "the internet have suffered incrementally heavy compressions. The RECCE model, as a result, turns out to be the winner with the highest fake scores and the number of correctly classified faces when facing real-life Deepfake suspects. 7.2 Deepfake Intelligence Considering real-life usages, accessible Deepfake detection models such as the ones being discussed", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1544 }, { "text": "Deepfake Detection: A Comprehensive Survey from the Reliability Perspective 25 regarding the Deepfake detection research domain by endlessly integrating the detection models and real-life Deepfake intelligence beyond detection. 8 DISCUSSION Our study provides a scientific workflow to prove the reliability of the detection models when", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1545 }, { "text": "applied to real-life cases. Unlike research that simply enumerates detection performance on each benchmark dataset, the reliability study scheme derives statistical claims regarding the detectors on arbitrary candidate Deepfake suspects with the help of confidence intervals. Particularly, the interval values are reliable based on sufficient trials from a sampling frame that ideally imitates", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1546 }, { "text": "real-world Deepfake distribution according to CLT. Considering that the prosecutors are unable to prove the fake-video claim due to the challenge brought to the video evidence authentication standard by Deepfake, the experiment results in Section 4 have solved the problem favorably. Specifically, the model reliability study scheme can be qualified by expert witnesses for the validity", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1547 }, { "text": "of the detection models based on the expert\u2019s testimony following corresponding rules, and thus, the reliable statistical metrics regarding the detection performance may assist the video evidence for criminal investigation cases. The accuracies are picked to assist the claims while the AUC scores are more helpful at the research level.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1548 }, { "text": "As a result, a reliable justification can be claimed based on values in Table 4 and Table 5 with the help of the confidence intervals and mean values once a sampling option is chosen. For example, suppose the RECCE model is used to help justify the authenticity of the cheerleader video of the \u2018Deepfake mom\u2019 case following the balanced sampling results, a claim can be made such that the", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1549 }, { "text": "video is fake with accuracies lying in the range of [60.37%, 61.22%] and [60.29%, 61.30%] with 90% and 95% confidence levels, respectively. In other words, we are 90% and 95% confident to declare that the video is fake with accuracies in the range of [60.37%, 61.22%] and [60.29%, 61.30%], respectively. If the imbalanced sampling option is trusted, a similar claim can be concluded in the", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1550 }, { "text": "range of [62.36%, 63.66%] and [62.24%, 63.79%] accuracies for the 90% and 95% confidence levels, respectively. In real life, since the authenticity of the candidate suspect is normally unknown, based on the experiment results in this study, the dominant MAT model is likely to be adopted for Deepfake", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1551 }, { "text": "detection and the following conclusion can be provided following the balanced sampling setting: we are 95% confident to classify the video as real (or fake) with an accuracy between 68.41% and 69.37%. If the imbalanced sampling setting is trusted, the winning Xception model can be employed to offer the justification such that the video is real (or fake) with an accuracy between 65.37% and", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1552 }, { "text": "67.07% at the 95% confidence level. Meanwhile, several findings can be concluded based on the results in Table 2 and Tables 4 to 8. Firstly, trade-offs are objectively unavoidable when attempting each of the three challenges. Specifically, regarding Table 2, models with promising transferability on lab-controlled benchmark", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1553 }, { "text": "datasets might lack interpretability on their performance and robustness in sophisticated real-life scenarios. While successfully explaining the detection decision with pellucid evidence and common sense or smoothly resolving challenges in specific real-life conditions with robustness as reported in the published papers, there usually remains limited computational power and model ability to", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1554 }, { "text": "enrich the transferability and detection accuracy on unseen data. Secondly, although the detection performance in Sections 6.2 and 6.3 is unremarkable compared to other approaches, the RECCE model appears to derive the best detection results on real-life Deepfake videos given the fact that we are aware that they are all fake, while on the contrary, the winning MAT and Xception models", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1555 }, { "text": "26 Wang et al. materials. This may be caused by the potential adversarial attacks such that the facial manipulation technique of the fake materials can easily fool models that are mainly based on certain feature extraction perspectives or techniques, and the detectors, therefore, need to be improved to better", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1556 }, { "text": "cooperate with the reliability study scheme. Lastly, for models that have a large gap between accuracy and AUC score values, it may be meaningful to locate a classification threshold other than 0.5 in order to achieve satisfactory detection accuracy since their high AUC scores have highlighted the ability to separate real and fake data.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1557 }, { "text": "9 CONCLUSION This paper provides a thorough survey of reliability-oriented Deepfake detection approaches by defining the three challenges of Deepfake detection research: transferability, interpretability, and robustness. While the early methods mainly solve puzzles on seen data, improvements by persistent", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1558 }, { "text": "attempts have gradually shown promising up-to-date detection performance on unseen benchmark datasets. Considering the lack of usage for the well-trained detection models to benefit real life and even specifically for criminal investigation, this paper conducts a comprehensive survey regarding model reliability by introducing an unprecedented model reliability study scheme that bridges", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1559 }, { "text": "the research gap and provides a reliable path for applying on-the-shelf detection models to assist prosecutions in court for Deepfake related cases. A barely discussed standardized data preparation workflow is simultaneously designed and presented for the reference of both starters and veterans in the domain. The reliable accuracies of the detection models derived by random sampling at the", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1560 }, { "text": "90% and 95% confidence levels are informative and may be adopted as or to assist forensic evidence in court for Deepfake-related cases under the qualification of expert testimony. Based on the informative findings in validating the detection models, potential future research trends are worth discussing. Although a Deepfake detection model can be verified the reliability", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1561 }, { "text": "for real-life usages via the presented reliability study scheme in this paper, an ideal model that simultaneously solves transferability, interpretability, and robustness challenges is not yet accom- plished in the current research domain. Consequently, obvious trade-offs have been observed when resolving each of the three challenges. Therefore, considering that the ground-truth attacks and", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1562 }, { "text": "labels of current and future Deepfake contents will not be visible to victims in Deepfake-related cases, researchers may continuously advance the detection model performance regarding each of the three challenges, but more importantly, a model that satisfies all three goals at the same time is urgently desired. At the same time, based on the model reliability study scheme that is first put", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1563 }, { "text": "forward in this study, subsequent improvements and discussions can also be conducted to achieve the general reliability goal progressively. For instance, tracing original sources of synthetic con- tents and recovering synthetic operation sequences are worth exploiting in future work to further enhance the reliability of a falsification. Moreover, although videos in this study are either real or", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1564 }, { "text": "fake as a whole, as the real-world Deepfake materials become complex, videos with only partial image frames being fake can lead to further potential risks, and the corresponding benchmark datasets and detectors are also desired. REFERENCES [1] U.S Food & Drug Administration. 2022. COVID-19 Antigen Home Test Package Insert for Healthcare Providers.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1565 }, { "text": "https://www.fda.gov/media/152698/download. Accessed: 2022-09-09. [2] Darius Afchar, Vincent Nozick, Junichi Yamagishi, and Isao Echizen. 2018. MesoNet: a Compact Facial Video Forgery Detection Network. 2018 IEEE International Workshop on Information Forensics and Security (WIFS) (2018), 1\u20137. [3] Shruti Agarwal, Hany Farid, Tarek El-Gaaly, and Ser-Nam Lim. 2020. Detecting Deep-Fake Videos from Appearance", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1566 }, { "text": "Deepfake Detection: A Comprehensive Survey from the Reliability Perspective 27 [4] Irene Amerini, Leonardo Galteri, Roberto Caldelli, and Alberto Del Bimbo. 2019. Deepfake Video Detection through Optical Flow Based CNN. In 2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW). IEEE,", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1567 }, { "text": "1205\u20131207. [5] Rima Sabina Aouf. 2019. Museum creates deepfake Salvador Dal\u00ed to greet visitors. https://www.dezeen.com/2019/05/ 24/salvador-dali-deepfake-dali-musuem-florida/. Accessed: 2022-09-08. [6] Yong Bai, Yuanfang Guo, Jinjie Wei, Lin Lu, Rui Wang, and Yunhong Wang. 2020. Fake Generated Painting Detection", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1568 }, { "text": "Via Frequency Analysis. In 2020 IEEE International Conference on Image Processing (ICIP). 1256\u20131260. https://doi.org/ 10.1109/ICIP40778.2020.9190892 [7] Federico Baldassarre, Quentin Debard, Gonzalo Fiz Pontiveros, and Tri Kurniawan Wijaya. 2022. Quantitative Metrics for Evaluating Explanations of Video DeepFake Detectors. In 33rd British Machine Vision Conference 2022, BMVC 2022,", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1569 }, { "text": "London, UK, November 21-24, 2022. BMVA Press. https://bmvc2022.mpi-inf.mpg.de/0972.pdf [8] Le Minh Binh and Simon Woo. 2022. ADD: Frequency Attention and Multi-View Based Knowledge Distillation to Detect Low-Quality Compressed Deepfake Images. Proceedings of the AAAI Conference on Artificial Intelligence 36, 1", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1570 }, { "text": "(Jun. 2022), 122\u2013130. https://doi.org/10.1609/aaai.v36i1.19886 [9] Nicol\u00f2 Bonettini, Edoardo Daniele Cannas, Sara Mandelli, Luca Bondi, Paolo Bestagini, and Stefano Tubaro. 2021. Video Face Manipulation Detection Through Ensemble of CNNs. In 2020 25th International Conference on Pattern Recognition (ICPR). 5012\u20135019. https://doi.org/10.1109/ICPR48806.2021.9412711", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1571 }, { "text": "a \"Siamese\" Time Delay Neural Network. In Proceedings of the 6th International Conference on Neural Information Processing Systems (NIPS\u201993). 737\u2013744. [12] R.S. Brown. 2010. Sampling. In International Encyclopedia of Education (Third Edition) (third edition ed.), Penelope Peterson, Eva Baker, and Barry McGaw (Eds.). Elsevier, Oxford, 142\u2013146. https://doi.org/10.1016/B978-0-08-044894-", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1572 }, { "text": "7.00294-3 [13] Junyi Cao, Chao Ma, Taiping Yao, Shen Chen, Shouhong Ding, and Xiaokang Yang. 2022. End-to-End Reconstruction- Classification Learning for Face Forgery Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 4113\u20134122. [14] Shenhao Cao, Qin Zou, Xiuqing Mao, Dengpan Ye, and Zhongyuan Wang. 2021. Metric Learning for Anti-Compression", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1573 }, { "text": "Facial Forgery Detection. In Proceedings of the 29th ACM International Conference on Multimedia (ACM MM 2021). 1929\u20131937. [15] Nicholas Carlini and Hany Farid. 2020. Evading Deepfake-Image Detectors With White- and Black-Box Attacks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1574 }, { "text": "Examples: Towards Good Generalizations for DeepFake Detections. In CVPR. [18] Liang Chen, Yong Zhang, Yibing Song, Jue Wang, and Lingqiao Liu. 2022. OST: Improving Generalization of DeepFake Detection via One-Shot Test-Time Training. In Advances in Neural Information Processing Systems, Alice H. Oh, Alekh", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1575 }, { "text": "Agarwal, Danielle Belgrave, and Kyunghyun Cho (Eds.). [19] Renwang Chen, Xuanhong Chen, Bingbing Ni, and Yanhao Ge. 2020. SimSwap: An Efficient Framework For High Fidelity Face Swapping. In MM \u201920: The 28th ACM International Conference on Multimedia. [20] Shen Chen, Taiping Yao, Yang Chen, Shouhong Ding, Jilin Li, and Rongrong Ji. 2021. Local Relation Learning for", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1576 }, { "text": "Face Forgery Detection. Proceedings of the AAAI Conference on Artificial Intelligence 35, 2 (May 2021), 1081\u20131088. https://ojs.aaai.org/index.php/AAAI/article/view/16193 [21] Harry Cheng, Yangyang Guo, Tianyi Wang, Qi Li, Xiaojun Chang, and Liqiang Nie. 2023. Voice-Face Homogeneity Tells Deepfake. ACM Transactions on Multimedia Computing, Communications, and Applications 20, 3, Article 76 (nov", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1577 }, { "text": "2023), 22 pages. https://doi.org/10.1145/3625231 [22] Rudy Chinchilla. 2021. Mom Made Deepfake Nudes of Daughter\u2019s Cheer Teammates to Harass Them: Po- lice. https://www.nbcphiladelphia.com/news/local/mom-made-deepfake-nudes-of-daughters-cheer-teammates-to- harass-them-police/2740906/. Accessed: 2022-09-08.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1578 }, { "text": "28 Wang et al. [25] Yang Chuming, Daniel Wu, and Ken Hong. 2021. Practical Deepfake Detection: Vulnerabilities in Global Contexts. In Responsible AI (RAT) - ICLR 2021 workshop. [26] Umur Aybars Ciftci, Ilke Demir, and Lijun Yin. 2020. FakeCatcher: Detection of Synthetic Portrait Videos using Biological Signals. IEEE Transactions on Pattern Analysis and Machine Intelligence (2020), 1\u20131. https://doi.org/10.1109/", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1579 }, { "text": "TPAMI.2020.3009287 [27] EU Health Security Committee. 2022. EU Common list of COVID-19 antigen tests. https://health.ec.europa.eu/system/ files/2022-07/covid-19_eu-common-list-antigen-tests_en.pdf. Accessed: 2022-09-09. [28] Davide Cozzolino, Diego Gragnaniello, and Luisa Verdoliva. 2014. Image forgery localization through the fusion of", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1580 }, { "text": "camera-based, feature-based and pixel-based techniques. In 2014 IEEE International Conference on Image Processing (ICIP). 5302\u20135306. https://doi.org/10.1109/ICIP.2014.7026073 [29] D. Cozzolino and L. Verdoliva. 2020. Noiseprint: A CNN-Based Camera Model Fingerprint. IEEE Transactions on Information Forensics and Security 15 (2020), 144\u2013159. https://doi.org/10.1109/TIFS.2019.2916364", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1581 }, { "text": "Manipulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). [32] Catherine de Weever and S. Wilczek. 2020. Deepfake detection through PRNU and logistic regression analyses. [33] deepfakes. 2018. FakeApp. https://www.malavida.com/en/soft/fakeapp/. Accessed: 2022-09-08.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1582 }, { "text": "Image Matching. In Computer Vision \u2013 ECCV 2022, Shai Avidan, Gabriel Brostow, Moustapha Ciss\u00e9, Giovanni Maria Farinella, and Tal Hassner (Eds.). Springer Nature Switzerland, Cham, 18\u201335. [40] Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1583 }, { "text": "Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In International Conference on Learning Representations. [41] Ricard Durall, Margret Keuper, Franz-Josef Pfreundt, and Janis Keuper. 2020. Unmasking DeepFakes with simple Features. https://doi.org/10.48550/ARXIV.1911.00686 [42] Hany Farid. 2022. Creating, Using, Misusing, and Detecting Deep Fakes. Journal of Online Trust and Safety 1, 4 (Sep.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1584 }, { "text": "2022). https://doi.org/10.54501/jots.v1i4.56 [43] Pasquale Ferrara, Tiziano Bianchi, Alessia De Rosa, and Alessandro Piva. 2012. Image Forgery Localization via Fine-Grained Analysis of CFA Artifacts. IEEE Transactions on Information Forensics and Security 7 (2012), 1566\u20131577. [44] FFmpeg. 2021. FFmpeg. https://www.ffmpeg.org/. Accessed: 2021-08-29.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1585 }, { "text": "Frequency Analysis for Deep Fake Image Recognition. In Proceedings of the 37th International Conference on Machine Learning (ICML\u201920). JMLR.org, Article 304, 12 pages. [47] Jessica Fridrich and Jan Kodovsky. 2012. Rich Models for Steganalysis of Digital Images. IEEE Transactions on Information Forensics and Security 7, 3 (2012), 868\u2013882. https://doi.org/10.1109/TIFS.2012.2190402", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1586 }, { "text": "for Identity Swapping. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 3403\u20133412. https://doi.org/10.1109/CVPR46437.2021.00341 [50] Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative Adversarial Nets. In The 27th Neural Information Processing Systems Advances.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1587 }, { "text": "Deepfake Detection: A Comprehensive Survey from the Reliability Perspective 29 [51] Dou Goodman, Hao Xin, Wang Yang, Wu Yuesheng, Xiong Junfeng, and Zhang Huan. 2020. Advbox: a toolbox to generate adversarial examples that fool neural networks. arXiv:2001.05574 [cs.LG] [52] Hong Kong Special Administrative Region Government. 2022. Rapid Antigen Test (RAT) for COVID-19. https:", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1588 }, { "text": "//www.coronavirus.gov.hk/pdf/RapAgTest_FAQ_ENG.pdf. Accessed: 2022-09-09. [53] Qiqi Gu, Shen Chen, Taiping Yao, Yang Chen, Shouhong Ding, and Ran Yi. 2022. Exploiting Fine-Grained Face Forgery Clues via Progressive Enhancement Learning. Proceedings of the AAAI Conference on Artificial Intelligence 36, 1 (Jun.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1589 }, { "text": "2022), 735\u2013743. https://doi.org/10.1609/aaai.v36i1.19954 [54] Jiazhi Guan, Hang Zhou, Zhibin Hong, Errui Ding, Jingdong Wang, Chengbin Quan, and Youjian Zhao. 2022. Delving into Sequential Patches for Deepfake Detection. In Advances in Neural Information Processing Systems, Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho (Eds.).", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1590 }, { "text": "Traces on Images. IEEE Access 8 (2020), 165085\u2013165098. https://doi.org/10.1109/ACCESS.2020.3023037 [57] Nick Dufourand Andrew Gully. 2019. Contributing Data to Deepfake Detection Research. https://ai.googleblog.com/ 2019/09/contributing-data-to-deepfake-detection.html. Accessed: 2022-09-08. [58] Zhiqing Guo, Gaobo Yang, Jiyou Chen, and Xingming Sun. 2021. Fake face detection via adaptive manipulation traces", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1591 }, { "text": "extraction network. Computer Vision and Image Understanding 204 (2021), 103170. https://doi.org/10.1016/j.cviu. 2021.103170 [59] Zhiqing Guo, Gaobo Yang, Jiyou Chen, and Xingming Sun. 2022. Exposing Deepfake Face Forgeries with Guided Residuals. https://doi.org/10.48550/ARXIV.2205.00753 [60] David G\u00fcera and Edward J. Delp. 2018. Deepfake Video Detection Using Recurrent Neural Networks. In 2018 15th", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1592 }, { "text": "IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS). 1\u20136. https://doi.org/10.1109/ AVSS.2018.8639163 [61] Alexandros Haliassos, Rodrigo Mira, Stavros Petridis, and Maja Pantic. 2022. Leveraging Real Talking Faces via Self- Supervision for Robust Forgery Detection. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1593 }, { "text": "(CVPR). 14930\u201314942. https://doi.org/10.1109/CVPR52688.2022.01453 [62] Drew Harwell. 2021. Remember the \u2018deepfake cheerleader mom\u2019? Prosecutors now admit they can\u2019t prove fake-video claims. https://www.washingtonpost.com/technology/2021/05/14/deepfake-cheer-mom-claims-dropped/. Accessed: 2022-09-08.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1594 }, { "text": "IEEE Conference on Computer Vision and Pattern Recognition. 770\u2013778. https://doi.org/10.1109/CVPR.2016.90 [65] Yang He, Ning Yu, Margret Keuper, and Mario Fritz. 2021. Beyond the Spectrum: Detecting Deepfakes via Re-synthesis. In 30th International Joint Conference on Artificial Intelligence (IJCAI).", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1595 }, { "text": "transformer. Applied Intelligence (2022). https://doi.org/10.1007/s10489-022-03867-9 [68] Sepp Hochreiter and J\u00fcrgen Schmidhuber. 1997. Long Short-term Memory. Neural computation 9 (12 1997), 1735\u201380. https://doi.org/10.1162/neco.1997.9.8.1735 [69] Ashish Hooda, Neal Mangaokar, Ryan Feng, Kassem Fawaz, Somesh Jha, and Atul Prakash. 2022. Towards Adversarially", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1596 }, { "text": "Robust Deepfake Detection: An Ensemble Approach. https://doi.org/10.48550/ARXIV.2202.05687 [70] Chih-Chung Hsu, Yi-Xiu Zhuang, and Chia-Yen Lee. 2020. Deep Fake Image Detection Based on Pairwise Learning. Applied Sciences 10, 1 (2020). https://doi.org/10.3390/app10010370 [71] Juan Hu, Xin Liao, Jinwen Liang, Wenbo Zhou, and Zheng Qin. 2022. FInfer: Frame Inference-Based Deepfake", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1597 }, { "text": "Detection for High-Visual-Quality Videos. Proceedings of the AAAI Conference on Artificial Intelligence 36, 1 (Jun. 2022), 951\u2013959. https://doi.org/10.1609/aaai.v36i1.19978 [72] Juan Hu, Xin Liao, Wei Wang, and Zheng Qin. 2022. Detecting Compressed Deepfake Videos in Social Networks Using Frame-Temporality Two-Stream Convolutional Network. IEEE Transactions on Circuits and Systems for Video", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1598 }, { "text": "Technology 32, 3 (2022), 1089\u20131102. https://doi.org/10.1109/TCSVT.2021.3074259 [73] Dong Huang and Fernando De La Torre. 2012. Facial Action Transfer with Personalized Bilinear Regression. In Computer Vision \u2013 ECCV 2012, Andrew Fitzgibbon, Svetlana Lazebnik, Pietro Perona, Yoichi Sato, and Cordelia Schmid", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1599 }, { "text": "30 Wang et al. [74] Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q. Weinberger. 2017. Densely Connected Convolutional Networks. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2261\u20132269. https://doi.org/10. 1109/CVPR.2017.243 [75] Shehzeen Hussain, Paarth Neekhara, Brian Dolhansky, Joanna Bitton, Cristian Canton Ferrer, Julian McAuley, and", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1600 }, { "text": "Farinaz Koushanfar. 2022. Exposing Vulnerabilities of Deepfake Detection Systems with Robust Attacks. Digital Threats 3, 3, Article 30 (feb 2022), 23 pages. https://doi.org/10.1145/3464307 [76] Wombo Studios Inc. 2021. Wombo: Make your selfies sing. https://play.google.com/store/apps/details?id=com.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1601 }, { "text": "womboai.wombo&hl=en&gl=US. Accessed: 2022-09-08. [77] Mousa Tayseer Jafar, Mohammad Ababneh, Mohammad Al-Zoube, and Ammar Elhassan. 2020. Forensics and Analysis of Deepfake Videos. In 2020 11th International Conference on Information and Communication Systems (ICICS). 053\u2013058. https://doi.org/10.1109/ICICS49469.2020.239493", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1602 }, { "text": "Using Frequency-Level Perturbations. Proceedings of the AAAI Conference on Artificial Intelligence 36, 1 (Jun. 2022), 1060\u20131068. https://doi.org/10.1609/aaai.v36i1.19990 [80] Shuai Jia, Chao Ma, Taiping Yao, Bangjie Yin, Shouhong Ding, and Xiaokang Yang. 2022. Exploring Frequency Adversarial Attacks for Face Forgery Detection. In IEEE/CVF Conference on Computer Vision and Pattern Recognition,", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1603 }, { "text": "CVPR 2022, New Orleans, LA, USA, June 18-24, 2022. IEEE, 4093\u20134102. https://doi.org/10.1109/CVPR52688.2022.00407 [81] Liming Jiang, Ren Li, Wayne Wu, Chen Qian, and Chen Change Loy. 2020. DeeperForensics-1.0: A Large-Scale Dataset for Real-World Face Forgery Detection. In CVPR. 2889\u20132898. [82] Tackhyun Jung, Sangwon Kim, and Keecheon Kim. 2020. DeepVision: Deepfakes Detection Using Human Eye Blinking", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1604 }, { "text": "Pattern. IEEE Access 8 (2020), 83144\u201383154. https://doi.org/10.1109/ACCESS.2020.2988660 [83] R. E. Kalman. 1960. A New Approach to Linear Filtering and Prediction Problems. Journal of Basic Engineering 82, 1 (03 1960), 35\u201345. https://doi.org/10.1115/1.3662552 [84] Wonjun Kang, Geonsu Lee, Hyung Il Koo, and Nam Ik Cho. 2022. One-Shot Face Reenactment on Megapixels.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1605 }, { "text": "https://doi.org/10.48550/ARXIV.2205.13368 [85] Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. 2018. Progressive Growing of GANs for Improved Quality, Stability, and Variation. In International Conference on Learning Representations. https://openreview.net/forum?id= Hk99zCeAb [86] T. Karras, S. Laine, and T. Aila. 2021. A Style-Based Generator Architecture for Generative Adversarial Networks.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1606 }, { "text": "IEEE Transactions on Pattern Analysis and Machine Intelligence 43, 12 (dec 2021), 4217\u20134228. https://doi.org/10.1109/ TPAMI.2020.2970919 [87] Katie Katro. 2022. Bucks County mother gets probation in harassment case involving daughter\u2019s cheerlead- ing rivals. https://6abc.com/raffaela-spone-bucks-county-pa-cheerleaders-harassment-case-victory-vipers-squad/", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1607 }, { "text": "11939419/. Accessed: 2022-09-08. [88] Leo Kelion. 2018. Deepfake porn videos deleted from internet by Gfycat. https://www.bbc.com/news/technology- 42905185. Accessed: 2022-09-08. [89] Jan Kietzmann, Linda W. Lee, Ian P. McCarthy, and Tim C. Kietzmann. 2020. Deepfakes: Trick or treat? Business Horizons 63, 2 (2020), 135\u2013146.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1608 }, { "text": "https://doi.org/10.1016/j.bushor.2019.11.006 ARTIFICIAL INTELLIGENCE AND MACHINE LEARNING. [90] Davis King. 2021. dlib 19.22.1. https://pypi.org/project/dlib/. Accessed: 2021-08-29. [91] Diederik P. Kingma and Max Welling. 2014. Auto-Encoding Variational Bayes. In 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14-16, 2014, Conference Track Proceedings, Yoshua Bengio", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1609 }, { "text": "and Yann LeCun (Eds.). [92] Marissa Koopman, Andrea Macarulla Rodriguez, and Zeno Geradts. 2018. Detection of Deepfake Video Manipulation. In Proceedings of the 20th Irish Machine Vision and Image Processing conference. 133\u2013136. [93] Pavel Korshunov and S\u00e9bastien Marcel. 2018. DeepFakes: a New Threat to Face Recognition? Assessment and", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1610 }, { "text": "Detection. CoRR abs/1812.08685 (2018). arXiv:1812.08685 http://arxiv.org/abs/1812.08685 [94] Prabhat Kumar, Mayank Vatsa, and Richa Singh. 2020. Detecting Face2Face Facial Reenactment in Videos. In 2020 IEEE Winter Conference on Applications of Computer Vision (WACV). 2578\u20132586. https://doi.org/10.1109/WACV45572.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1611 }, { "text": "Deepfake Detection: A Comprehensive Survey from the Reliability Perspective 31 [96] Laan Labs. 2017. Face Swap Live. https://play.google.com/store/apps/details?id=com.laan.labs.faceswaplive&hl=en& gl=US. Accessed: 2022-09-28. [97] Dave Lee. 2018. Deepfakes porn has serious consequences. https://www.bbc.com/news/technology-42912529.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1612 }, { "text": "Accessed: 2022-09-08. [98] Victor Lempitsky, Andrea Vedaldi, and Dmitry Ulyanov. 2018. Deep Image Prior. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 9446\u20139454. https://doi.org/10.1109/CVPR.2018.00984 [99] Jiaming Li, Hongtao Xie, Jiahong Li, Zhongyuan Wang, and Yongdong Zhang. 2021. Frequency-Aware Discriminative", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1613 }, { "text": "Feature Learning Supervised by Single-Center Loss for Face Forgery Detection. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021. Computer Vision Foundation / IEEE, 6458\u20136467. https://doi.org/10.1109/CVPR46437.2021.00639 [100] Lingzhi Li, Jianmin Bao, Hao Yang, Dong Chen, and Fang Wen. 2019. FaceShifter: Towards High Fidelity And", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1614 }, { "text": "Occlusion Aware Face Swapping. arXiv preprint arXiv:1912.13457 (2019). [101] Lingzhi Li, Jianmin Bao, Ting Zhang, Hao Yang, Dong Chen, Fang Wen, and Baining Guo. 2020. Face X-Ray for More General Face Forgery Detection. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 5000\u20135009. https://doi.org/10.1109/CVPR42600.2020.00505", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1615 }, { "text": "on Computer Vision and Pattern Recognition Workshops (CVPRW). [104] Yuezun Li, Xin Yang, Pu Sun, Honggang Qi, and Siwei Lyu. 2020. Celeb-DF: A Large-Scale Challenging Dataset for DeepFake Forensics. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 3204\u20133213. https://doi.org/10.1109/CVPR42600.2020.00327", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1616 }, { "text": "772\u2013781. https://doi.org/10.1109/CVPR46437.2021.00083 [108] Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV).", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1617 }, { "text": "en&gl=US. Accessed: 2022-09-28. [111] Shao-An Lu. 2018. faceswap-GAN. https://github.com/shaoanlu/faceswap-GAN. Accessed: 2022-09-08. [112] J. Lukas, J. Fridrich, and M. Goljan. 2006. Digital camera identification from sensor pattern noise. IEEE Transactions on Information Forensics and Security 1, 2 (2006), 205\u2013214. https://doi.org/10.1109/TIFS.2006.873602", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1618 }, { "text": "A Large-Scale Study. Journal of Imaging 7, 10 (2021). https://doi.org/10.3390/jimaging7100193 [115] Marek MarekKowalski. 2019. FaceSwap. https://github.com/MarekKowalski/FaceSwap. Accessed: 2021-08-29. [116] Francesco Marra, Giovanni Poggi, Carlo Sansone, and Luisa Verdoliva. 2017. Blind PRNU-Based Image Clustering", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1619 }, { "text": "for Source Identification. IEEE Transactions on Information Forensics and Security 12, 9 (2017), 2197\u20132211. https: //doi.org/10.1109/TIFS.2017.2701335 [117] Luca Martino, David Luengo, and Joaqu\u00edn M\u00edguez. 2018. Direct Methods. Springer International Publishing, Cham, 27\u201363. https://doi.org/10.1007/978-3-319-72634-2_2", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1620 }, { "text": "32 Wang et al. [119] Falko Matern, Christian Riess, and Marc Stamminger. 2019. Exploiting Visual Artifacts to Expose Deepfakes and Face Manipulations. In 2019 IEEE Winter Applications of Computer Vision Workshops (WACVW). 83\u201392. https: //doi.org/10.1109/WACVW.2019.00020 [120] Joshua Rhett Miller. 2022. Deepfake video of Zelensky telling Ukrainians to surrender removed from social plat-", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1621 }, { "text": "forms. https://nypost.com/2022/03/17/deepfake-video-shows-volodymyr-zelensky-telling-ukrainians-to-surrender/. Accessed: 2022-09-08. [121] Yisroel Mirsky and Wenke Lee. 2021. The Creation and Detection of Deepfakes: A Survey. ACM Comput. Surv. 54, 1, Article 7 (jan 2021), 41 pages. https://doi.org/10.1145/3425780", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1622 }, { "text": "(SIGGRAPH \u201918). Association for Computing Machinery, New York, NY, USA, Article 69, 2 pages. https://doi.org/10. 1145/3230744.3230818 [124] Huy Hoang Nguyen, Fuming Fang, Junichi Yamagishi, and Isao Echizen. 2019. Multi-task Learning for Detecting and Segmenting Manipulated Facial Images and Videos. 2019 IEEE 10th International Conference on Biometrics Theory,", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1623 }, { "text": "Applications and Systems (BTAS) (2019), 1\u20138. [125] Huy H. Nguyen, Junichi Yamagishi, and Isao Echizen. 2019. Use of a Capsule Network to Detect Fake Images and Videos. arXiv:1910.12467 [cs.CV] [126] Sophie J. Nightingale, Shruti Agarwal, Erik H\u00e4rk\u00f6nen, Jaakko Lehtinen, and Hany Farid. 2021. Synthetic faces: how", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1624 }, { "text": "perceptually convincing are they? Journal of Vision 21, 9 (2021), 2015. https://doi.org/10.1167/jov.21.9.2015 [127] Sophie J. Nightingale and Hany Farid. 2022. AI-synthesized faces are indistinguishable from real faces and more trustworthy. Proceedings of the National Academy of Sciences 119, 8 (2022), e2120481119. https://doi.org/10.1073/pnas.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1625 }, { "text": "2120481119 [128] Sophie J. Nightingale and Hany Farid. 2022. Synthetic Faces Are More Trustworthy Than Real Faces. Proceedings of the National Academy of Sciences 22, 14 (2022), 3068. https://doi.org/10.1167/jov.22.14.3068 [129] Yuval Nirkin, Yosi Keller, and Tal Hassner. 2019. FSGAN: Subject agnostic face swapping and reenactment. In", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1626 }, { "text": "Proceedings of the IEEE International Conference on Computer Vision. 7184\u20137193. [130] Yuval Nirkin, Lior Wolf, Yosi Keller, and Tal Hassner. 2022. DeepFake Detection Based on Discrepancies Between Faces and Their Context. IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 10 (2022), 6111\u20136121.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1627 }, { "text": "https://doi.org/10.1109/TPAMI.2021.3093446 [131] World Health Organization. 2022. Use of SARS-CoV-2 antigen-detection rapid diagnostic tests for COVID-19 self- testing. https://apps.who.int/iris/bitstream/handle/10665/352350/WHO-2019-nCoV-Ag-RDTs-Self-testing-2022.1- eng.pdf?sequence=1. Accessed: 2022-09-09.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1628 }, { "text": "tests to detect SARS-CoV-2 variants of concern. Medical Microbiology and Immunology 210, 5 (01 Dec 2021), 263\u2013275. https://doi.org/10.1007/s00430-021-00719-0 [133] Xunyu Pan, Xing Zhang, and Siwei Lyu. 2012. Exposing image splicing with inconsistent local noise variances. In 2012 IEEE International Conference on Computational Photography (ICCP). 1\u201310. https://doi.org/10.1109/ICCPhot.2012.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1629 }, { "text": "6215223 [134] Taesung Park, Ming-Yu Liu, Ting-Chun Wang, and Jun-Yan Zhu. 2019. Semantic Image Synthesis With Spatially- Adaptive Normalization. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2332\u20132341. https://doi.org/10.1109/CVPR.2019.00244 [135] Bo Peng, Wei Wang, Jing Dong, and Tieniu Tan. 2017. Optimized 3D Lighting Environment Estimation for Image", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1630 }, { "text": "Forgery Detection. IEEE Transactions on Information Forensics and Security 12 (2017), 479\u2013494. [136] Z. Peng, W. Huang, S. Gu, L. Xie, Y. Wang, J. Jiao, and Q. Ye. 2021. Conformer: Local Features Coupling Global Representations for Visual Recognition. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV). IEEE", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1631 }, { "text": "Computer Society, Los Alamitos, CA, USA, 357\u2013366. https://doi.org/10.1109/ICCV48922.2021.00042 [137] Ivan Perov, Daiheng Gao, Nikolay Chervoniy, Kunlin Liu, Sugasa Marangonda, Chris Um\u00e9, Mr. Dpfks, Carl Shift Facenheim, Luis RP, Jian Jiang, Sheng Zhang, Pingyu Wu, Bo Zhou, and Weiming Zhang. 2021. DeepFaceLab:", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1632 }, { "text": "Integrated, flexible and extensible face-swapping framework. arXiv:2005.05535 [cs.CV] [138] Francesco Picetti, Sara Mandelli, Paolo Bestagini, Vincenzo Lipari, and Stefano Tubaro. 2020. DIPPAS: A Deep Image Prior PRNU Anonymization Scheme. arXiv:2012.03581 [cs.MM] [139] Adam Polyak, Lior Wolf, and Yaniv Taigman. 2019. TTS Skins: Speaker Conversion via ASR.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1633 }, { "text": "Deepfake Detection: A Comprehensive Survey from the Reliability Perspective 33 [140] Hua Qi, Qing Guo, Felix Juefei-Xu, Xiaofei Xie, Lei Ma, Wei Feng, Yang Liu, and Jianjun Zhao. 2020. DeepRhythm: Exposing DeepFakes with Attentional Visual Heartbeat Rhythms. In Proceedings of the 28th ACM International", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1634 }, { "text": "Conference on Multimedia (Seattle, WA, USA) (MM \u201920). Association for Computing Machinery, New York, NY, USA, 4318\u20134327. https://doi.org/10.1145/3394171.3413707 [141] Yuyang Qian, Guojun Yin, Lu Sheng, Zixuan Chen, and Jing Shao. 2020. Thinking in Frequency: Face Forgery Detection by Mining Frequency-Aware Clues. In Computer Vision \u2013 ECCV 2020, Andrea Vedaldi, Horst Bischof,", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1635 }, { "text": "Thomas Brox, and Jan-Michael Frahm (Eds.). Springer International Publishing, Cham, 86\u2013103. [142] E. Reinhard, M. Adhikhmin, B. Gooch, and P. Shirley. 2001. Color transfer between images. IEEE Computer Graphics and Applications 21, 5 (2001), 34\u201341. https://doi.org/10.1109/38.946629 [143] Learn & revise. 2019. Deepfakes: What are they and why would I make one? https://www.bbc.co.uk/bitesize/articles/", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1636 }, { "text": "zfkwcqt. Accessed: 2022-09-08. [144] Andreas Rossler, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies, and Matthias Niessner. 2019. FaceForensics++: Learning to Detect Manipulated Facial Images. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 1\u201311.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1637 }, { "text": "the 31st International Conference on Neural Information Processing Systems (Long Beach, California, USA) (NIPS\u201917). Curran Associates Inc., Red Hook, NY, USA, 3859\u20133869. [147] Shota Saito, Yoichi Tomioka, and Hitoshi Kitazawa. 2017. A Theoretical Framework for Estimating False Acceptance Rate of PRNU-Based Camera Identification. IEEE Transactions on Information Forensics and Security 12, 9 (2017),", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1638 }, { "text": "2026\u20132035. https://doi.org/10.1109/TIFS.2017.2692683 [148] ScienceDaily. 2020. \u2018Deepfakes\u2019 ranked as most serious AI crime threat. https://www.sciencedaily.com/releases/2020/ 08/200804085908.htm. Accessed: 2021-05-01. [149] Shaikh Akib Shahriyar and Matthew Wright. 2022. Evaluating Robustness of Sequence-Based Deepfake Detector", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1639 }, { "text": "Models by Adversarial Perturbation. In Proceedings of the 1st Workshop on Security Implications of Deepfakes and Cheapfakes (WDC \u201922). Association for Computing Machinery, New York, NY, USA, 13\u201318. https://doi.org/10.1145/ 3494109.3527194 [150] Zhihua Shang, Hongtao Xie, Zhengjun Zha, Lingyun Yu, Yan Li, and Yongdong Zhang. 2021. PRRNet: Pixel-Region", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1640 }, { "text": "relation network for face forgery detection. Pattern Recognition 116 (2021), 107950. https://doi.org/10.1016/j.patcog. 2021.107950 [151] Kaede Shiohara and Toshihiko Yamasaki. 2022. Detecting Deepfakes with Self-Blended Images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 18720\u201318729.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1641 }, { "text": "Detection through Precise Geometric Features. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 3609\u20133618. [157] Mingxing Tan and Quoc Le. 2019. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. In Proceedings of the 36th International Conference on Machine Learning (Proceedings of Machine Learning Research,", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1642 }, { "text": "Vol. 97). PMLR, 6105\u20136114. [158] Shahroz Tariq, Sangyup Lee, Hoyoung Kim, Youjin Shin, and Simon S. Woo. 2018. Detecting Both Machine and Human Created Fake Face Images In the Wild. In Proceedings of the 2nd International Workshop on Multimedia Privacy and Security (Toronto, Canada) (MPS \u201918). Association for Computing Machinery, New York, NY, USA, 81\u201387.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1643 }, { "text": "https://doi.org/10.1145/3267357.3267367 [159] Justus Thies, Michael Zollh\u00f6fer, and Matthias Nie\u00dfner. 2019. Deferred Neural Rendering: Image Synthesis Using Neural Textures. ACM Trans. Graph. 38, 4, Article 66 (July 2019), 12 pages. [160] Justus Thies, Michael Zollhofer, Marc Stamminger, Christian Theobalt, and Matthias Niessner. 2016. Face2Face: Real-", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1644 }, { "text": "Time Face Capture and Reenactment of RGB Videos. In IEEE Conference on Computer Vision and Patten Recognition (CVPR). 2387\u20132395. [161] Global Times. 2021. Chinese social media platforms delete actor\u2019s accounts of for hurting the nation after controversial photos of Yasukuni Shrine. https://www.globaltimes.cn/page/202108/1231473.shtml. Accessed: 2022-09-27.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1645 }, { "text": "34 Wang et al. [162] Global Times. 2021. Chinese surrogacy scandal actress Zheng Shuang fined $46 million for tax evasion, shows banned. https://www.globaltimes.cn/page/202108/1232636.shtml. Accessed: 2022-09-27. [163] Global Times. 2021. Works of scandals-hit actress Zhao Wei removed from platforms, following ban on actor Zhang", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1646 }, { "text": "Zhehan for visiting Yasukuni Shrine. https://www.globaltimes.cn/page/202108/1232631.shtml. Accessed: 2022-09-27. [164] Ruben Tolosana, Sergio Romero-Tapiador, Julian Fierrez, and Ruben Vera-Rodriguez. 2021. DeepFakes Evolution: Analysis of Facial Regions and Fake Detection Performance. In Pattern Recognition. ICPR International Workshops", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1647 }, { "text": "and Challenges, Alberto Del Bimbo, Rita Cucchiara, Stan Sclaroff, Giovanni Maria Farinella, Tao Mei, Marco Bertini, Hugo Jair Escalante, and Roberto Vezzani (Eds.). Springer International Publishing, Cham, 442\u2013456. [165] Ruben Tolosana, Ruben Vera-Rodriguez, Julian Fierrez, Aythami Morales, and Javier Ortega-Garcia. 2020. Deepfakes", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1648 }, { "text": "and beyond: A Survey of face manipulation and fake detection. Information Fusion 64 (2020), 131\u2013148. https: //doi.org/10.1016/j.inffus.2020.06.014 [166] David Tonucci. 2005. 44 - New and Emerging Testing Technology for Efficacy and Safety E valuation of Personal Care Delivery Systems. In Delivery System Handbook for Personal Care and Cosmetic Products, Meyer R. Rosen (Ed.).", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1649 }, { "text": "William Andrew Publishing, Norwich, NY, 911\u2013929. https://doi.org/10.1016/B978-081551504-3.50049-3 [167] Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou. 2021. Training data-efficient image transformers & distillation through attention. In Proceedings of the 38th International", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1650 }, { "text": "Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 139), Marina Meila and Tong Zhang (Eds.). PMLR, 10347\u201310357. [168] Loc Trinh, Michael Tsang, Sirisha Rambhatla, and Yan Liu. 2021. Interpretable and Trustworthy Deepfake Detection via Dynamic Prototypes. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV).", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1651 }, { "text": "1973\u20131983. [169] Soumya Tripathy, Juho Kannala, and Esa Rahtu. 2019. ICface: Interpretable and Controllable Face Reenactment Using GANs. arXiv preprint arXiv:1904.01909 (2019). [170] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Advances in Neural Information Processing Systems, I. Guyon, U. V.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1652 }, { "text": "Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.), Vol. 30. Curran Associates, Inc. [171] Sheng-Yu Wang, Oliver Wang, Richard Zhang, Andrew Owens, and Alexei A. Efros. 2020. CNN-Generated Images Are Surprisingly Easy to Spot... for Now. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1653 }, { "text": "(CVPR). 8692\u20138701. https://doi.org/10.1109/CVPR42600.2020.00872 [172] Tianyi Wang, Harry Cheng, Kam Pui Chow, and Liqiang Nie. 2023. Deep Convolutional Pooling Transformer for Deepfake Detection. ACM Transactions on Multimedia Computing, Communications, and Applications 19, 6. [173] Tianyi Wang and Kam Pui Chow. 2023. Noise Based Deepfake Detection via Multi-Head Relative-Interaction.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1654 }, { "text": "Proceedings of the AAAI Conference on Artificial Intelligence (2023). [174] Tianyi Wang, Ming Liu, Wei Cao, and Kam Pui Chow. 2022. Deepfake noise investigation and detection. Forensic Science International: Digital Investigation 42 (2022), 301395. https://doi.org/10.1016/j.fsidi.2022.301395 Proceedings", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1655 }, { "text": "of the Twenty-Second Annual DFRWS USA. [175] Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. 2021. Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction Without Convolutions. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 568\u2013578.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1656 }, { "text": "Conferences on Artificial Intelligence Organization, 1136\u20131142. https://doi.org/10.24963/ijcai.2021/157 [177] Yukai Wang, Chunlei Peng, Decheng Liu, Nannan Wang, and Xinbo Gao. 2022. ForgeryNIR: Deep Face Forgery and Detection in Near-Infrared Scenario. IEEE Transactions on Information Forensics and Security 17 (2022), 500\u2013515.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1657 }, { "text": "https://doi.org/10.1109/TIFS.2022.3146766 [178] Mika Westerlund. 2019. The Emergence of Deepfake Technology: A Review. Technology Innovation Management Review 9 (11 2019), 40\u201353. https://doi.org/10.22215/timreview/1282 [179] Deressa Wodajo and Solomon Atnafu. 2021. Deepfake Video Detection Using Convolutional Vision Transformer.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1658 }, { "text": "https://arxiv.org/abs/2102.11126 [180] Walt Woods, Jack Chen, and Christof Teuscher. 2019. Adversarial explanations for understanding image classification decisions and improved neural network robustness. Nature Machine Intelligence 1, 11 (01 Nov 2019), 508\u2013516. https://doi.org/10.1038/s42256-019-0104-6", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1659 }, { "text": "Deepfake Detection: A Comprehensive Survey from the Reliability Perspective 35 [182] Haiwei Wu, Jiantao Zhou, Jinyu Tian, Jun Liu, and Yu Qiao. 2022. Robust Image Forgery Detection Against Trans- mission Over Online Social Networks. IEEE Transactions on Information Forensics and Security 17 (2022), 443\u2013456.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1660 }, { "text": "https://doi.org/10.1109/TIFS.2022.3144878 [183] Wayne Wu, Yunxuan Zhang, Cheng Li, Chen Qian, and Chen Change Loy. 2018. ReenactGAN: Learning to Reenact Faces via Boundary Transfer. In ECCV. [184] Xi Wu, Zhen Xie, YuTao Gao, and Yu Xiao. 2020. SSTNet: Detecting Manipulated Faces Through Spatial, Steganalysis", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1661 }, { "text": "and Temporal Features. In ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 2952\u20132956. https://doi.org/10.1109/ICASSP40776.2020.9053969 [185] Ying Xu, Kiran Raja, and Marius Pedersen. 2022. Supervised Contrastive Learning for Generalizable and Explainable", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1662 }, { "text": "DeepFakes Detection. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) Workshops. 379\u2013389. [186] Xin Yang, Yuezun Li, and Siwei Lyu. 2019. Exposing Deep Fakes Using Inconsistent Head Poses. In ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 8261\u20138265. https://doi.org/10.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1663 }, { "text": "1109/ICASSP.2019.8683164 [187] Egor Zakharov, Aliaksandra Shysheya, Egor Burkov, and Victor Lempitsky. 2019. Few-Shot Adversarial Learning of Realistic Neural Talking Head Models. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV). 9458\u20139467. https://doi.org/10.1109/ICCV.2019.00955", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1664 }, { "text": "learning of deep CNN for image denoising. IEEE Transactions on Image Processing 26, 7 (2017), 3142\u20133155. [190] Shanghang Zhang, Xiaohui Shen, Zhe Lin, Radom\u00edr Mech, Jo\u00e3o P. Costeira, and Jose M.F. Moura. 2018. Learning to Understand Image Blur. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 6586\u20136595.", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1665 }, { "text": "https://doi.org/10.1109/CVPR.2018.00689 [191] Yunxuan Zhang, Siwei Zhang, Yue He, Cheng Li, Chen Change Loy, and Ziwei Liu. 2019. One-shot Face Reenactment. In British Machine Vision Conference (BMVC). [192] Hanqing Zhao, Wenbo Zhou, Dongdong Chen, Tianyi Wei, Weiming Zhang, and Nenghai Yu. 2021. Multi-Attentional", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1666 }, { "text": "Deepfake Detection. In IEEE Conference on Computer Vision and Patten Recognition (CVPR). 2185\u20132194. [193] T. Zhao, X. Xu, M. Xu, H. Ding, Y. Xiong, and W. Xia. 2021. Learning Self-Consistency for Deepfake Detection. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV). IEEE Computer Society, Los Alamitos, CA, USA,", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1667 }, { "text": "15003\u201315013. https://doi.org/10.1109/ICCV48922.2021.01475 [194] Peng Zhou, Xintong Han, Vlad I. Morariu, and Larry S. Davis. 2017. Two-Stream Neural Networks for Tampered Face Detection. In 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). 1831\u20131839. https://doi.org/10.1109/CVPRW.2017.229", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1668 }, { "text": "Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR). 4834\u20134844. [197] Bojia Zi, Minghao Chang, Jingjing Chen, Xingjun Ma, and Yu-Gang Jiang. 2020. WildDeepfake: A Challenging Real-World Dataset for Deepfake Detection. 2382\u20132390. https://doi.org/10.1145/3394171.3413769", "source": "Deepfake Detection Reliability Survey", "year": 2022, "url": "https://arxiv.org/abs/2211.10881", "id": 1669 }, { "text": "The Tug-of-War Between Deepfake Generation and Detection Hannah Lee 1 2 Changyeon Lee 3 4 5 Kevin Farhat 1 2 Lin Qiu 1 2 Steve Geluso 2 Aerin Kim 5 2 Oren Etzioni 2 Abstract Multimodal generative models are rapidly evolv- ing, leading to a surge in the generation of realistic video and audio that offers exciting possibilities", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1670 }, { "text": "but also serious risks. Deepfake videos, which can convincingly impersonate individuals, have particularly garnered attention due to their po- tential misuse in spreading misinformation and creating fraudulent content. This survey paper examines the dual landscape of deepfake video generation and detection, emphasizing the need", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1671 }, { "text": "for effective countermeasures against potential abuses. We provide a comprehensive overview of current deepfake generation techniques, including face swapping, reenactment, and audio-driven ani- mation, which leverage cutting-edge technologies like generative adversarial networks and diffusion models to produce highly realistic fake videos.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1672 }, { "text": "Additionally, we analyze various detection ap- proaches designed to differentiate authentic from altered videos, from detecting visual artifacts to deploying advanced algorithms that pinpoint in- consistencies across video and audio signals. The effectiveness of these detection methods heav- ily relies on the diversity and quality of datasets", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1673 }, { "text": "used for training and evaluation. We discuss the evolution of deepfake datasets, highlighting the importance of robust, diverse, and frequently up- dated collections to enhance the detection accu- racy and generalizability. As deepfakes become increasingly indistinguishable from authentic con- tent, developing advanced detection techniques", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1674 }, { "text": "that can keep pace with generation technologies is crucial. We advocate for a proactive approach in 1Paul G. Allen School of Computer Science & Engineering, University of Washington, Seattle WA, USA 2TrueMedia.org, Seat- tle WA, USA 3Department of Computer Science and Engineering, Yonsei University, Seoul, Republic of Korea 4Computer and Infor-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1675 }, { "text": "mation Technology, Purdue University, West Lafayette IN, USA 5Miraflow, Kirkland WA, USA. Correspondence to: Hannah Lee . Data-centric Machine Learning Research Workshop (DMLR) at the CNN 41 st International Conference on Machine Learning, Vienna, Austria. PMLR 235, 2024. Copyright 2024 by the author(s).", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1676 }, { "text": "the \u201ctug-of-war\u201d between deepfake creators and detectors, emphasizing the need for continuous research collaboration, standardization of evalua- tion metrics, and the creation of comprehensive benchmarks. 1. Introduction Recent advances in multimodal generative models have made manipulated media increasingly more realistic and", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1677 }, { "text": "accessible. Although synthetically generated audio, images, and videos can have creative and beneficial applications such as improved dubbing or translation of films (Yang et al., 2020; Hu et al., 2021), deepfake videos that impersonate humans highlight the potential harms of media manipula- tion and synthetic generation. For example, deepfakes that", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1678 }, { "text": "blend celebrities\u2019 faces onto bodies in pornographic videos (Nguyen et al., 2022; Hsu, 2024) and alter politicians\u2019 mes- sages (Suwajanakorn et al., 2017) can spread misinforma- tion, threaten individuals, and damage reputations (Zhou & Zafarani, 2020), disrupting election campaigns and financial markets. Recently, deepfakes have also become integral to", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1679 }, { "text": "fraudulent schemes, resulting in scams of up to $25 mil- lion (Chen & Magramo, 2024). Given the rise of social media and online media consumption, it is unsurprising that deepfake videos are increasingly interfering with people\u2019s lives. The first modern deepfake videos surfaced in 2017 when users on Reddit posted computer-generated pornographic", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1680 }, { "text": "videos of actresses (Nguyen et al., 2022). Since then, many deepfake generation tools have become available for pub- lic use. FaceSwap (Fac, 2024), FaceSwapGAN (Lu, 2024), StyleGAN (Karras et al., 2019), FSGAN (Nirkin et al., 2019) and other applications allow anyone with basic program- ming skills to generate their own deepfakes, and as text-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1681 }, { "text": "to-video models such as Imagen Video (Ho et al., 2022), CogVideo (Hong et al., 2022), and Sora (Brooks et al., 2024) improve and become widespread, the barrier to entry will continue to be lowered. As deepfake video generation has become increasingly de- mocratized and capable of producing realistic results, em-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1682 }, { "text": "The Tug-of-War Between Deepfake Generation and Detection follows two patterns: detection and prevention. While pre- vention through techniques like watermarking (Lv, 2021) and blockchain frameworks (Rashid et al., 2021) as well as through technology policy (Reisach, 2021) is critical for mitigating deepfake harms, it is out of scope for this paper", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1683 }, { "text": "and we instead focus on deepfake detection, the currently more common strategy. Deepfake video detection is comprised of fake image de- tection, fake audio detection, and techniques specific to video sequences. These techniques are generally leveraged for the binary classification task where classifiers learn to", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1684 }, { "text": "distinguish authentic and manipulated videos. Initial de- tection tools focused primarily on visual artifacts such as blended face edges (Li et al., 2020a) and image forgery techniques that have preceded deepfakes (Chaitra & Reddy, 2022). More recently, deep learning techniques have capi- talized on the large amounts of real and fake data available", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1685 }, { "text": "online. However, since training detection algorithms de- pends on fake data created by generation tools, deepfake detectors lag behind generators. Consequently, the devel- opment of detection algorithms provides direct feedback to generation algorithms on what makes deepfakes detectable and can encourage adversarial generation to bypass detec-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1686 }, { "text": "tion. We refer to this relationship as the ongoing tug-of-war between deepfake generation and detection. There are existing survey papers that explore deepfake gen- eration and detection (Rana et al., 2022; Yu et al., 2021; Yi et al., 2023b; Mirsky & Lee, 2021; Nguyen et al., 2022; Swathi & Sk, 2021). However, these often only focus on de-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1687 }, { "text": "tection (Rana et al., 2022; Yu et al., 2021; Yi et al., 2023b); do not include modern methods that have gained popular- ity (Mirsky & Lee, 2021; Swathi & Sk, 2021); or do not consider in depth the impact of the datasets used to train generation and detection models. In our approach, we focus on video generation and detection while discussing the cur-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1688 }, { "text": "rent landscape of curating relevant training and evaluation datasets. Our main contributions include an updated survey of deepfake generation and detection techniques; identifica- tion of the successes, challenges, and limitations of current deepfake detection practices; and suggestions for future deepfake detection research, focusing on the importance of", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1689 }, { "text": "quality datasets in this tug-of-war. 2. Deepfake Video Generation Generating deepfake video media consists of generating both visual and audio content. Here, we briefly describe common generation processes for both modalities. 2.1. Face Swapping Face swapping is one method used to create deepfake im-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1690 }, { "text": "ages and videos by replacing one person\u2019s face with an- other\u2019s face. This technique is now widely available through well-developed packages like Face Fusion (Ruhs, 2024) and Faceswap (Fac, 2024). Early works could only handle faces in the same pose, but later developments have incorporated 3D-based methods to construct faces with variations (Dale", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1691 }, { "text": "et al., 2011; Lin et al., 2012). With the development of neural networks, face swapping can now involve an encoder- decoder network (Perov et al., 2020). The encoder extracts latent features from the faces in the source image while the decoder reconstructs the target image using these features. More recently, GAN-based methods have also been applied;", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1692 }, { "text": "FSGAN (Nirkin et al., 2019) and FSGAN2 (Nirkin et al., 2022) provide a two-stage pipeline that supports both face swapping and reenactment (further discussed in Section 2.2) simultaneously. These methods train a generator to learn the latent representation of the source image and then inpaint the segmented area in the target image. These GAN-based", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1693 }, { "text": "methods are able to generate more realistic images with higher resolution and fidelity. However, training GANs can be unstable, restricting their actual application to lower resolution images (Xu et al., 2022b; Liu et al., 2023b). 2.2. Reenactment Reenactment focuses on making one person\u2019s face in a video", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1694 }, { "text": "mimic the facial expressions and head movements of an- other person. Unlike face swapping, the identity of the face remains the same while its expressions and movements are altered (Nguyen et al., 2022). Reenactment can drive the expression, gaze, mouth, pose, identity, and body of the target subject based on the source subject (Mirsky & Lee,", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1695 }, { "text": "2021). Face2Face (Thies et al., 2016) proposed the first method that achieves real-time RGB-only reenactment with- out the need for a teeth proxy (Garrido et al., 2016; Thies et al., 2015) or direct source-to-target copying (Vlasic et al., 2005). FSGAN can be applied to unseen pairs of faces and adjusts significant pose and expression variations that can", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1696 }, { "text": "be applied to a single image or a video sequence (Nirkin et al., 2019). Later, FSGAN2 extended the model by intro- ducing preprocessing and additional post-processing steps for reducing flickering and saturation artifacts (Nirkin et al., 2022). Recently, facial reenactment efforts have not only focused on updating the expressions on the target image", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1697 }, { "text": "itself but have also involved a joint effort in creating both audio and visual components for talking face generation (Section 2.4). 2.3. Diffusion-based Deepfakes Diffusion-based deepfake generation represents a signifi- cant advancement over traditional methods like GANs and autoencoders, offering more realistic and efficient capabil-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1698 }, { "text": "The Tug-of-War Between Deepfake Generation and Detection refined this approach. DDPM utilizes a series of learned noise adjustments to reverse the addition of noise, effec- tively reconstructing the data distribution of original images. Subsequently, the Latent Diffusion Model (LDM) (Rom- bach et al., 2021) enhances this process by generating in the", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1699 }, { "text": "latent space, speeding up the diffusion process and handling complex data distributions. Image Generation. Diffusion-generated images are cate- gorized into two types: text-to-image, where users provide text prompts to specify image content to models such as DALL-E (Ramesh et al., 2021), DreamBooth (Ruiz et al.,", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1700 }, { "text": "2022) and Imagen (Saharia et al., 2022); and text guided im- age editing or generation, which transforms existing images based on text prompts using Imagic (Kawar et al., 2022), InstructPix2Pix (Brooks et al., 2022), or other tools. Chen et al. (2023) proposes a method to specifically generate high", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1701 }, { "text": "quality deepfakes by leveraging diffusion, using celebrity images as a guide for the image\u2019s latent initialization. Video Generation. DMs are also capable of generating videos. Typically, this is done by inflating 2D convolutional layers into pseudo-3D convolutional layers in the DM and incorporating temporal self-attention layers in each trans-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1702 }, { "text": "former block to manage the temporal consistency across frames. As one example, Diffusion Heads (Stypu\u0142kowski et al., 2024) leverages an autoregressive DM that takes one identity image and an audio sequence to generate a talking head, which can also be applied to the deepfake generation with a custom identity image.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1703 }, { "text": "2.4. Audio-Driven Facial Animation The field of audio-driven facial animation has been a fas- cinating area of exploration for the computer vision and graphics community. Many studies have been carried out on both digital 3D human faces (Karras et al., 2017a; Richard et al., 2021; Zhou et al., 2018; Fan et al., 2022) and realistic", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1704 }, { "text": "human head generation (Fan et al., 2015; Suwajanakorn et al., 2017; Jamaludin et al., 2019). Talking Head Generation. Early research on lip-syncing mostly focused on creating the entire head of a speaking person (Suwajanakorn et al., 2017; Jamaludin et al., 2019; Wang et al., 2020a; Zhang et al., 2023). Several methods", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1705 }, { "text": "used structural details like 2D (Chen et al., 2019) landmarks, 3D landmarks (Zhou et al., 2020) and 3D meshes (Chen et al., 2020). However, the resulting output often exhibits inconsistencies between connecting parts and suffers from a lack of precision. Many methods that are specific to in- dividuals (Suwajanakorn et al., 2017; Lu et al., 2021; Ji", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1706 }, { "text": "et al., 2021) depend on the subject and may struggle with generalization. Furthermore, the recent application of NeRF for person-specific modeling, as cited in several works (Guo et al., 2021; Liu et al., 2022; Shen et al., 2022), often exhibits sub-optimal performance, including low visual quality, re-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1707 }, { "text": "stricted expressiveness, and inconsistency across frames, when the available training data is limited. Lip-Syncing on Faces. Other research has focused on lip- syncing the mouth in videos while leaving other elements unchanged (Prajwal et al., 2020; Park et al., 2022; Thies et al., 2020). Specifically, Wav2Lip (Prajwal et al., 2020) is", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1708 }, { "text": "capable of generating lip-sync results that are not person- specific. However, the generative model used is built on low- resolution images, resulting in blurry outputs. It also falls short in capturing the unique identity of a target including the shape of teeth and lips when provided with a target", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1709 }, { "text": "template video. 2.5. Style-based Generator for Faces StyleGAN models (Karras et al., 2019; 2020; 2021) have shown great success on image generation tasks, particularly on facial image generation and editing (Abdal et al., 2019; 2020; Tov et al., 2021; Richardson et al., 2021). They have also been leveraged for face restorations (Wang et al., 2021;", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1710 }, { "text": "Yang et al., 2021) and face swapping (Xu et al., 2022c;a), which require the preservation of the original facial emo- tion and expressions. Concurrently, style-based generators have been effectively utilized in the creation of facial ani- mations (Burkov et al., 2020; Liang et al., 2022). However, the generators rely on the style vectors in the W or W + la-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1711 }, { "text": "tent spaces for controlling both the appearances and motion dynamics. A significant limitation of this approach is the W + space\u2019s inability to maintain the spatial consistency of backgrounds. This often results in non-realistic outcomes or noticeable artifacts. 2.6. Audio Generation Deepfake audio generation involves creating realistic syn-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1712 }, { "text": "thetic audio, often mimicking a specific person\u2019s voice. We describe common techniques used to produce highly con- vincing audio deepfakes. Text-to-Speech (TTS). Text-to-speech systems convert writ- ten text into spoken words. Modern TTS systems use deep learning models to produce natural-sounding speech. Key", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1713 }, { "text": "approaches include: \u2022 Concatenative TTS: A traditional method involving concatenating pre-recorded speech segments to gener- ate fluent speech. While it can produce high-quality and natural-sounding results by leveraging real audio, it often struggles with capturing the dynamic variations in human speech (Hande, 2014). This can lead to no-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1714 }, { "text": "The Tug-of-War Between Deepfake Generation and Detection set of parameters that describe aspects of the speech signal such as pitch, duration, and spectral features by using statistical models like Hidden Markov Mod- els (Tokuda et al., 2000). This approach offers flexibil- ity in creating various voices and speaking styles, but", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1715 }, { "text": "the synthesized speech may sound less natural due to the oversimplification of speech variations. \u2022 Neural TTS: Using deep learning models to generate more natural and expressive speech (Tan et al., 2021). Examples include Tacotron 2 (Shen et al., 2017) and WaveNet (van den Oord et al., 2016), where Tacotron 2", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1716 }, { "text": "converts text to a mel-spectrogram then uses a modified WaveNet as a vocoder to generate audio waveforms from the spectrogram. Voice Conversion (VC). VC aims to modify a source speaker\u2019s voice to sound like a target speaker without chang- ing the linguistic content. This process typically involves extracting features from the source speech, such as pitch,", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1717 }, { "text": "spectral envelope, and timbre, then transforming these fea- tures to match the target speaker\u2019s characteristics. Statistical models like Gaussian Mixture Models (GMMs) (Kain & Macon, 1998) and deep learning approaches such as GANs and variational autoencoders (VAEs) are employed to learn the mapping between the source and target features. Recent", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1718 }, { "text": "advancements have significantly improved the naturalness and intelligibility of converted speech (Sisman et al., 2020). Emotion Fake. Emotion fake, or emotional VC, involves altering the emotional tone of a speaker\u2019s voice. This tech- nique is essential for creating more expressive and engaging synthetic speech. It typically involves extracting prosodic", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1719 }, { "text": "features such as pitch, duration, and intensity before modify- ing these features to reflect the desired emotion. Deep learn- ing methods, particularly Convolutional Neural Networks (CNNs), GANs and Recurrent Neural Networks (RNNs), have shown promising results in this area (Trinh Van et al., 2022; Liu et al., 2020a). These models can learn complex", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1720 }, { "text": "mappings between neutral and emotional speech, resulting in more natural and convincing emotional expressions. Scene Fake. Scene fake, also known as environmental sound synthesis, involves generating background sounds or environmental effects to accompany synthetic speech to enhance the overall experience. This technique aims to", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1721 }, { "text": "create a more realistic auditory scene by simulating various ambient sounds such as street noise, office chatter, or natural environments like piano or birdsong (Donahue et al., 2019; Huzaifah & Wyse, 2020). GANs and autoencoder-based architectures have been applied to synthesize scene fake audio. Partially Fake. Partially altered audio involves modifying", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1722 }, { "text": "specific words or segments within a spoken recording by replacing them with either authentic or artificially gener- ated audio snippets. The speaker\u2019s voice remains consistent throughout the original speech and the altered segments, making the modified audio sound authentic despite the in- troduced changes. This technique allows for counterfeit", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1723 }, { "text": "versions of the initial speech to be created while maintain- ing the speaker\u2019s identity and vocal characteristics (Wu et al., 2022). 3. Deepfake Video Detection Deepfake video generation techniques necessitate a broad arsenal of detection tools. Fake video detection is gener- ally broken down into three main categories; fake image", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1724 }, { "text": "detection where individual frames are analyzed; fake audio detection to analyze the audio of a video; and fake video detection, which may utilize both images and audio as well as temporal data. 3.1. Fake Image Detection Detecting fake images precedes deepfakes and deep learn- ing. We discuss methods ranging from artifact detection to", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1725 }, { "text": "modern methods that have developed with advancements in deep learning. General Visual Artifact Detection. Deepfake images may introduce subtle artifacts that are not present in real images. These can include artifacts introduced by face blending or face warping when swapping faces, as well as inconsisten-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1726 }, { "text": "cies in the overall image. For example, Li et al. (2020a) propose the \u201cface X-ray\u201d image representation to detect anomalies in the blending boundaries of faces in blended images. Other work has focused on the differing image tex- tures between generated and real content (Liu et al., 2020b); inconsistencies in head positions (Yang et al., 2019); the", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1727 }, { "text": "resolution variability that arises from warping faces (Li & Lyu, 2018); and missing reflections or details in the teeth and eyes (Matern et al., 2019). Many other visual artifact detection methods have been proposed, but as deepfake generation has improved and fewer artifacts remain, these techniques have become overshadowed by methods focused", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1728 }, { "text": "on detecting more subtle identifiers. Detecting GAN-generated Images. A subset of image deepfake detection methods have focused on detecting ar- tifacts unique to popular GAN models. Some detection methods still rely on visible differences such as irregular pupil shapes (Guo et al., 2022), while many others exploit", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1729 }, { "text": "lower level abnormalities. For example, the upsampling operations involved in GAN generation introduce model specific artifacts into the images\u2019 spatial and frequency do- mains (Zhang et al., 2019; Marra et al., 2019; Yu et al., 2021). Wang et al. (2020b) trained a ResNet-50 model (He et al., 2016) that detects these artifacts and found that GAN-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1730 }, { "text": "The Tug-of-War Between Deepfake Generation and Detection ter (Wang et al., 2019) can effectively detect AI-synthesized fake faces generated by GANs. Building upon these detec- tors, PatchForensics (Chai et al., 2020) introduced a detector that analyzes smaller patches of images to determine if there", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1731 }, { "text": "are AI-generated or manipulated areas. Diffusion Detection. Generalization of GAN-based detec- tion methods to newer, diffusion based image generation techniques is difficult (Corvi et al., 2023; Ojha et al., 2023). The artifacts introduced by GANs in images are no longer present with DMs. Recently, Wang et al. (2023) proposed", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1732 }, { "text": "DIRE, a new image representation that uses reconstruc- tions of images using DMs as a method for detecting DM generated images. They hypothesize that DM generated images consist of features that are better reconstructed by other pretrained DMs, compared to reconstructions of real images. Furthermore, Lim et al. (2024) introduced Dis-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1733 }, { "text": "tilDIRE, a diffusion-generated image detection framework that significantly reduces the computational demands of the original DIRE method. Ojha et al. (2023) instead utilize the learned feature space of a pretrained vision-language model to determine if an image was AI-generated. And Lorenz et al. (2023) rely on multi Local Intrinsic Dimensionality to", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1734 }, { "text": "detect diffusion. These novel methods highlight how detec- tion algorithms continue to adapt to the newer generative models. 3.2. Fake Audio Detection Audio deepfake detection can utilize traditional classifiers like Support Vector Machines (SVM) after feature extrac- tion using methods such as STFT spectrograms or Mel-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1735 }, { "text": "frequency Cepstral coefficients (MFCC). However, deep learning models outperform traditional classifiers by learn- ing complex patterns, resulting in better accuracy; these approaches are now widely preferred (Zaman et al., 2023). Typically, deep learning-based audio deepfake detection involves extracting image features like STFT spectrogram", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1736 }, { "text": "feature images, followed by employing a CNN-based ar- chitecture to extract embeddings. A binary classification layer is then added to classify as real or fake. However, other techniques that are often employed in traditional audio classification can also be applied. RNN-based models can capture temporal patterns of audio signals due to their se-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1737 }, { "text": "quential nature (Zaman et al., 2023), enabling classification into different categories (Scarpiniti et al., 2021; Gimeno et al., 2020). More recently, transformer-based models have been intro- duced. These models, when compared to CNN-based ones, can handle input-length variance due to their multi-head", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1738 }, { "text": "self-attention mechanisms (Zaman et al., 2023). This allows them to effectively capture useful global-context informa- tion, regardless of the audio length (Gong et al., 2021; Ghosh et al., 2023; Liu et al., 2023a). 3.3. Fake Video Detection Deepfake detection techniques for videos transcend detec-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1739 }, { "text": "tion techniques for single images by comparing images temporally frame after frame. Deepfake video detection techniques may also combine information from images and audio to make inferences based on coherence, synchroniza- tion, and physiology. Frame by Frame Analysis. Evidence of deepfake video generation may be apparent analyzing videos frame by", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1740 }, { "text": "frame. Determining movement via optical flow may dis- tinguish real videos from fake videos (Amerini et al., 2019). Guera and Delp (2018) use a CNN and RNN to extract fea- tures in frames and detect inconsistencies from the overall frame sequence, acknowledging deepfake manipulations often only appear briefly in overall videos. Taking one", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1741 }, { "text": "frame, and predicting the next frame, and measuring the error between the prediction and the actual next frame pro- vides a basis for prediction error energy analysis (Amerini & Caldelli, 2020). Frame analysis techniques can be improved by pre-processing video data to focus on area where faces occur (Sabir et al., 2019).", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1742 }, { "text": "Recent frame by frame methods include GenCon- ViT (Wodajo & Atnafu, 2023), which generalizes deep- fake video detection by extracting latent spaces of video frames using two networks with an independently trained autoencoder and VAE. The VAE reconstructs images and compares the reconstruction to the sample image. The au-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1743 }, { "text": "toencoder and VAE each feed into a ConvNeXt layer, and a Swin Transformer forming a hybrid model ConvNeXt-Swin which learns the relationships among the latent features. Sun et al. (2023) also analyzes frames, tracking the horizon- tal and vertical displacement trajectories of virtual anchor points on faces over time. They found real videos produce", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1744 }, { "text": "smoother trajectories than anchor points on fake faces and developed a fake trajectory detection network to classify real and fake videos. Physiological Features. Generative techniques often lack artifacts of biological processes. Deepfake video detection techniques have been able to identify deepfakes by analyz-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1745 }, { "text": "ing how people in the video blink (Li et al., 2018), whether the shape of their mouth is in synchronization with the sounds their mouth is making (Agarwal et al., 2020), and measuring heart rate via photoplethysmography (PPG) sig- nals from hemoglobin content in the blood changing how skin reflects light (Umur Aybars Ciftci, 2020).", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1746 }, { "text": "Audio-visual Analysis. Some work has observed incon- sistencies between images and audio. Chugh and Subra- manian (2020) observe that deepfake videos often have dissonance between the audio and video and create a metric called the Modality Dissonance Score. Computing audio- visual dissimilarity over 1-second video clips, they then", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1747 }, { "text": "The Tug-of-War Between Deepfake Generation and Detection aggregate these scores to perform classification. 3.4. Adversarial Attacks and Evading Detectors Despite numerous detection methods, studies have found that evasion efforts can be effective at rendering some detec- tors useless. Carlini and Farid (2020) found multiple small", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1748 }, { "text": "perturbation attacks that can be applied to deepfake images, resulting in the performance of the classifier by Wang et al. (2020b) to be reduced to worse than chance while min- imizing distortions visible to humans. Hou et al. (2023) and Neekhara et al. (2021) build on this work to extend the attacks and apply to other detectors. While these works do", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1749 }, { "text": "not include evasions of more modern detection methods, they speak to the tug-of-war nature between generation and detection. 4. Detection Challenges, Competitions, and Datasets Each of the methods described in Sections 2 and 3 have been trained on datasets of curated real and fake images. Here, we discuss the current landscape of available datasets", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1750 }, { "text": "and the progression of modern deepfake video detection algorithms. Table 1 summarizes some commonly used, publicly available datasets. 4.1. Deepfake Image Datasets Many datasets of real faces already exist for facial recogni- tion tasks. In 2007, the Labeled Faces in the Wild (LFW) database curated over 13,000 images of faces from the in-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1751 }, { "text": "ternet (Huang et al., 2007). Though now retracted, the MS- Celeb-1M dataset consisted of one million real images of 100,000 different identities (Guo et al., 2016). And IMDB- WIKI (Rothe et al., 2018) and IMDB-Clean (Lin et al., 2021) include 524,230 and 287,683 images respectively of faces, designed to train age-estimation algorithms. These datasets", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1752 }, { "text": "also aid the training of face image generation models. For deepfake image generation techniques that rely on face swapping or reenactment, the deepfake creator must have sufficient training data focused on both subjects (Fac, 2024; Ruhs, 2024). This has limited deepfakes to be of celebrities with large amounts of publicly available content, but as the", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1753 }, { "text": "quantity of public data increases and generation techniques improve, the accessibility to creating deepfakes continues to rise. The CelebA-HQ dataset (Karras et al., 2017b) is one early dataset of 30,000 images of the faces of celebrities, based on the larger CelebA dataset (Liu et al., 2015), that", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1754 }, { "text": "has been used to train GANs to generate higher resolution images of human faces. Karras et al. (2019) went on to create Flickr-Faces-HQ (FFHQ), a larger and more diverse dataset of 70,000 images no longer limited to celebrities, inclusive of broader ranges of ages, ethnicities, and image backgrounds.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1755 }, { "text": "In contrast to deepfake generation models, detection mod- els are typically trained on large datasets of both real and fake images. Often, these are custom datasets depending on the detection focus and currently available image gen- erators. Wang et al. (2020b) trained an image classifier on the LSUN (Yu et al., 2015) dataset and ProGAN (Karras", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1756 }, { "text": "et al., 2019) generated images. To test their image classifier, Wang et al. generated over 72,000 images using 11 differ- ent CNN-based models including StyleGAN (Karras et al., 2019), BigGAN (Brock et al., 2018), StarGAN (Choi et al., 2018), and CycleGAN (Zhu et al., 2017). This collection of real and fake samples has been used by subsequent de-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1757 }, { "text": "tection methods to train and evaluate models (Ojha et al., 2023; Corvi et al., 2023), but since it lacks newer genera- tion methods, new custom evaluations appear often (Wang et al., 2023; Lorenz et al., 2023; Lu & Ebrahimi, 2024). Standalone efforts have also been made to curate synthetic image datasets (Wang et al., 2022; Bird & Lotfi, 2024),", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1758 }, { "text": "which can be used to train detectors. However, while there have been attempts to standardize deepfake video detec- tion benchmarks (see Section 4.3), there are no commonly established evaluations for manipulated image detection. 4.2. Deepfake Audio Datasets and Detection Challenges As synthetic audio becomes increasingly accessible, interest", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1759 }, { "text": "in synthetic speech detection has grown (Yi et al., 2023b; Al- mutairi & Elgibreen, 2022; Hamza et al., 2022). This trend has spurred the creation of competitions like the ASVspoof Challenge (Wang et al., 2020c; Yamagishi et al., 2021), which aims to reduce reliance on specific knowledge of speaker verification systems and spoofing attacks, encourag-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1760 }, { "text": "ing a more realistic evaluation scenario reflecting real-world examples, and the Audio Deepfake Detection (ADD) Chal- lenge (Yi et al., 2022; 2023a), which aims to identify and analyze deepfake speech utterances, tackling the growing threats from advancements in speech synthesis and VC tech- nologies.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1761 }, { "text": "Beyond standarized challenges, research on audio deepfake detection has been energized by the introduction of various datasets. For example, the \u201cIn-the-Wild\u201d dataset (M\u00a8uller et al., 2022) contains audio deepfakes featuring politi- cians and public figures gathered from the internet. The ASVspoof DF dataset (Wang et al., 2020c; Yamagishi et al.,", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1762 }, { "text": "2021) comprises genuine and spoofed speech recordings that have been subjected to various lossy codecs such as m4a that are commonly employed in media storage. The DEEP- VOICE dataset (Bird & Lotfi, 2023) curates authentic human speech recordings from eight notable individuals, along with voices altered to mimic each other through Retrieval-based", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1763 }, { "text": "The Tug-of-War Between Deepfake Generation and Detection Table 1. Summary of common datasets used for training and evaluating deepfake detection models. Dataset Modality Identities Real Samples Generated Samples Generation Methods Year CNNDetect Image / 72,400 72,400 Multiple CNNs 2020 CIFAKE Image", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1764 }, { "text": "/ 60,000 60,000 Diffusion 2024 FoR Audio / >111,000 >87,000 TTS 2019 ASVspoof (LA) Audio 107 12,483 108,978 TTS, VC 2019 H-Voice Audio / 3,268 3,404 Multiple 2020 WaveFake Audio / / 117,985 Multiple 2021 In-the-Wild Audio 58 20.7 hours 17.2 hours Multiple 2022 EmoFake Audio 10 17,500 36,400 Multiple EVC Models", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1765 }, { "text": "2022 SceneFake Audio 107 19,838 64,642 Multiple 2022 DEEP-VOICE Audio 8 62 min 22 sec 62 min 22 sec RVC model 2023 ADD Audio / 243,194 273,874 Multiple 2023 DeepfakeTIMIT Video 32 320 640 GAN (face swap) 2018 FaceForensics++ Video / 1,000 4,000 Multiple 2019 Celeb-DF Video 59 590 5,639 GAN (face swap)", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1766 }, { "text": "2019 WildDeepfake Video 707 3,805 3,509 Multiple 2020 DFDC Video 960 23,654 104,500 Multiple 2020 DeeperForensics-1.0 Video 100 50,000 10,000 DF-VAE (face swap) 2020 AV-Deepfake1M Video 2,068 286,721 860,039 Multiple 2023 & Sch\u00a8onherr, 2021), EmoFake (Zhao et al., 2023), Scene- Fake (Yi et al., 2024), and H-Voice (Ballesteros et al., 2020)", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1767 }, { "text": "offer specialized perspectives, each contributing uniquely to the field. The FoR Dataset comprises over 198,000 ut- terances from deep-learning speech synthesizers and real speech, serving as a cornerstone for research in speech syn- thesis and synthetic speech detection. WaveFake provides samples from various network architectures and languages;", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1768 }, { "text": "EmoFake explores the impact of altering audio emotions; and SceneFake aims to detect manipulated audio scenes through modifications of the acoustic environment. Lastly, H-Voice consists of 6,672 histograms derived from authentic and synthetic voice recordings. 4.3. Deepfake Video Datasets and Detection Challenges", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1769 }, { "text": "One of the earliest and most influential datasets in the do- main of deepfake video detection is the FaceForensics++ dataset (R\u00a8ossler et al., 2019). It includes over 1,000 real video sequences collected from YouTube and corresponding deepfakes created using four different manipulation meth- ods: Deepfakes (Fac, 2024), Face2Face (Thies et al., 2016),", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1770 }, { "text": "FaceSwap (Marek, 2024), and NeuralTextures (Thies et al., 2019). The dataset is widely used for training and bench- marking detection algorithms due to its diversity in manip- ulation techniques, but remains limited by the small num- ber of unique identities present in the videos. Other early works include the WildDeepfake dataset (Zi et al., 2020),", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1771 }, { "text": "which contains a set of 7314 sequences of deepfake con- tent collected from the internet, and the DeepfakeTIMIT dataset (Korshunov & Marcel, 2018), a collection of public real videos each with a GAN generated deepfake counter- part. Building on the foundational work of datasets like Face- Forensics++, the DeepFake Detection Challenge (DFDC)", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1772 }, { "text": "dataset was introduced by Facebook AI to spur advance- ments in detection technologies (Dolhansky et al., 2020). The DFDC dataset is one of the most comprehensive public datasets, containing over 100,000 face swap video clips of both real and manipulated content. It is notable for having faces of more than 3,000 subjects and using 8 facial mod-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1773 }, { "text": "ification algorithms. It uses several deepfake generation techniques including GAN-based and non-learned methods to produce the clips. Another significant contribution is the Celeb-DF dataset (Li et al., 2020b), which addresses some of the limitations found in earlier datasets such as visual artifacts. Celeb-DF con-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1774 }, { "text": "tains 590 YouTube videos of celebrities of varying gender, age, and ethnicity, and 5,639 high-quality deepfake videos of these subjects generated using improved deepfake algo- rithms that reduce common visual artifacts, thereby posing a greater challenge for detection systems. In addition to these datasets, other efforts have focused on", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1775 }, { "text": "The Tug-of-War Between Deepfake Generation and Detection 11,000 manipulated deepfake videos generated using Deep- Fake VAE swapping. This dataset is subjected to a wide range of real-world distortions, resulting in a larger and more diverse collection of face swap videos that better re- flect the complexities of real-world scenarios.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1776 }, { "text": "AV-Deepfake1M (Cai et al., 2023) is one of the largest deepfake datasets to date, containing over 2,000 subjects and 1 million videos with various types of manipulations, including video, audio, and audio-visual deepfakes. Before this dataset, there had been limited datasets including small segments of audio-visual manipulations embedded within", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1777 }, { "text": "real videos, leaving detectors susceptible to these types of attacks. 5. Discussion The development of deepfake technologies has significantly advanced, leading to a continuous tug-of-war between gen- eration techniques and the corresponding detection methods. This dynamic interplay shapes the landscape of both fields.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1778 }, { "text": "5.1. Challenges in Current Detection Approaches Data Scarcity and Bias. One major issue is the lack of com- prehensive and diverse datasets that reflect the full range of manipulations generation models can produce. This scarcity leads to detection models that may perform well on specific types of deepfakes but fail to generalize to out-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1779 }, { "text": "of-distribution media, including new or slightly different deepfake approaches. Detection datasets are often custom-made, focusing on out- puts from specific generation models. This approach allows detectors to identify unique artifacts but limits their abil- ity to generalize. Additionally, relying on custom datasets", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1780 }, { "text": "complicates the direct comparison of detection approaches. Recent efforts such as DeepfakeBench (Yan et al., 2023) have made strides towards standardizing comparisons be- tween models by creating a benchmark incorporating many different datasets and streamlining the evaluation of models through a series of analysis tools. However, the rapid ad-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1781 }, { "text": "vancements in generators make detection datasets become quickly outdated. Evolving Generation Techniques. As generation meth- ods evolve, they often develop capabilities to bypass spe- cific detection mechanisms, especially those relying on de- tecting artifacts or inconsistencies that newer models no", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1782 }, { "text": "longer produce: the frequency space artifacts characteris- tic of GAN-generated images are noticeably absent from diffusion-generated images (Corvi et al., 2023) and adver- sarial perturbations can evade detectors (Carlini & Farid, 2020). Relatedly, the visibly increased resolution and re- alism of deepfakes make it challenging for human experts", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1783 }, { "text": "and automated systems to distinguish between genuine and manipulated content. This raises concerns about the effec- tiveness of current detection technologies and how to best curate datasets if the authenticity of media cannot be easily verified. 5.2. Future Directions Curate Robust Datasets and Design Competitions for", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1784 }, { "text": "Detection. To effectively combat deepfakes, there is a criti- cal need for creating and maintaining robust, diverse, and representative datasets that are publicly available to the re- search community. These datasets should include a wide variety of deepfake types and techniques to better train and", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1785 }, { "text": "test detection models. Diversity in modality should also be present; these datasets should include different types of media such as videos, audio, and images from various demo- graphics and in multiple languages to ensure comprehensive coverage. The datasets should be continuously updated at a regular cadence to include the latest deepfake techniques", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1786 }, { "text": "and real-world examples, ensuring that detection models can learn to counter new threats. By updating datasets in this way, deepfake detection methods can be trained more closely to detect the in-the-wild examples that pose the greatest threats to mislead individuals. Creating competi- tions for model evaluations can also establish standardized", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1787 }, { "text": "comparisons that test detectors on these pertinent examples as well. Focused Efforts on Representation and Consent. The use of web-scraped data to train both generation and detec- tion models also raises concerns about representation and consent. Deepfake subjects are often public figures like celebrities since ample data is available for them online.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1788 }, { "text": "However, as the barrier to create deepfakes lowers, issues around non-consensual deepfake pornography and identity misuse will likely become more prevalent. While such con- tent should be moderated and prevented from being created through policy and moderation considerations, detection methods must also be proactive in handling such cases.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1789 }, { "text": "Harness Capabilities of Foundation Models for Deep- fake Detection. One avenue parallel to curating robust datasets for training detection models is to harness the capa- bilities of foundation models that have been pretrained on large amounts of data, likely already inclusive of deepfake datasets. Some detection methods have started to utilize", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1790 }, { "text": "The Tug-of-War Between Deepfake Generation and Detection 6. Conclusion We provide a short survey of the current landscape of deepfake video generation and detection, as well as of the datasets used to train and evaluate these methods. Although great progress has been made in deepfake video detection,", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1791 }, { "text": "there is still much room for improvement in curating datasets that enable the development of robust, generalizable, and responsible detectors. Closer collaboration between the generation and detection communities could help anticipate future directions and ensure that detectors are well-equipped to handle the latest deepfake techniques. Regular detection", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1792 }, { "text": "competitions using shared and frequently updated datasets are a promising avenue to align research efforts to minimize deepfake misinformation spread and related consequences. References Faceswap, May 2024. URL https://github.com/ deepfakes/faceswap. original-date: 2017-12- 19T09:44:13Z. Abdal, R., Qin, Y., and Wonka, P. Image2stylegan: How", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1793 }, { "text": "to embed images into the stylegan latent space? In Pro- ceedings of the IEEE/CVF International Conference on Computer Vision, pp. 4432\u20134441, 2019. Abdal, R., Qin, Y., and Wonka, P. Image2stylegan++: How to edit the embedded images? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1794 }, { "text": "Recognition, pp. 8296\u20138305, 2020. Agarwal, S., Farid, H., Fried, O., and Agrawala, M. Detect- ing deep-fake videos from phoneme-viseme mismatches. pp. 660\u2013661, 2020. Almutairi, Z. and Elgibreen, H. A review of modern au- dio deepfake detection methods: Challenges and future directions. Algorithms, 15(5):155, 2022.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1795 }, { "text": "Amerini, I. and Caldelli, R. Exploiting prediction error inconsistencies through lstm-based classifiers to detect deepfake videos. IH&MMSec 2020, 2020. Amerini, I., Galteri, L., Caldelli, R., and Del Bimbo, A. Deepfake video detection through optical flow based cnn. pp. 1205\u20131207, 2019. doi: 10.1109/ICCVW.2019.00152.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1796 }, { "text": "Ballesteros, D. M., Rodriguez, Y., and Renza, D. A dataset of histograms of original and fake voice recordings (h-voice). Data in Brief, 29:105331, 2020. ISSN 2352- 3409. doi: https://doi.org/10.1016/j.dib.2020.105331. URL https://www.sciencedirect.com/ science/article/pii/S2352340920302250. Bird, J. J. and Lotfi, A. Real-time detection of ai-generated", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1797 }, { "text": "speech for deepfake voice conversion, 2023. Bird, J. J. and Lotfi, A. Cifake: Image classification and ex- plainable identification of ai-generated synthetic images. IEEE Access, 2024. Brock, A., Donahue, J., and Simonyan, K. Large scale gan training for high fidelity natural image synthesis. arXiv", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1798 }, { "text": "preprint arXiv:1809.11096, 2018. Brooks, T., Holynski, A., and Efros, A. A. Instruct- pix2pix: Learning to follow image editing instruc- tions. 2023 IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition (CVPR), pp. 18392\u201318402, 2022. URL https://api.semanticscholar. org/CorpusID:253581213.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1799 }, { "text": "Brooks, T., Peebles, B., Holmes, C., DePue, W., Guo, Y., Jing, L., Schnurr, D., Taylor, J., Luhman, T., Luhman, E., Ng, C., Wang, R., and Ramesh, A. Video generation models as world simulators. 2024. URL https://openai.com/research/ video-generation-models-as-world-simulators. Burkov, E., Pasechnik, I., Grigorev, A., and Lempitsky, V.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1800 }, { "text": "Neural head reenactment with latent pose descriptors. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 13786\u201313795, 2020. Cai, Z., Ghosh, S., Adatia, A. P., Hayat, M., Dhall, A., and Stefanov, K. Av-deepfake1m: A large-scale llm-driven audio-visual deepfake dataset, 2023.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1801 }, { "text": "Carlini, N. and Farid, H. Evading deepfake-image detectors with white-and black-box attacks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, pp. 658\u2013659, 2020. Chai, L., Bau, D., Lim, S.-N., and Isola, P. What makes fake images detectable? understanding properties that", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1802 }, { "text": "generalize, 2020. Chaitra, B. and Reddy, P. B. Digital image forgery: taxonomy, techniques, and tools\u2013a com- prehensive study. International Journal of System Assurance Engineering and Management, 14:18\u201333, 2022. URL https://api.semanticscholar. org/CorpusID:255039112. Chen, H. and Magramo, K. Finance worker pays out $25", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1803 }, { "text": "million after video call with deepfake \u2018chief financial officer\u2019, 2024. Chen, L., Maddox, R. K., Duan, Z., and Xu, C. Hierarchical cross-modal talking face generation with dynamic pixel- wise loss. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7832\u2013 7841, 2019.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1804 }, { "text": "The Tug-of-War Between Deepfake Generation and Detection Chen, L., Cui, G., Kou, Z., Zheng, H., and Xu, C. What comprises a good talking-head video generation?: A sur- vey and benchmark. arXiv preprint arXiv:2005.03201, 2020. Chen, Y., Haldar, N. A. H., Akhtar, N., and Mian, A. Text-image guided diffusion model for gener-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1805 }, { "text": "ating deepfake celebrity interactions. 2023 Inter- national Conference on Digital Image Computing: Techniques and Applications (DICTA), pp. 348\u2013355, 2023. URL https://api.semanticscholar. org/CorpusID:262825503. Choi, Y., Choi, M., Kim, M., Ha, J.-W., Kim, S., and Choo, J. Stargan: Unified generative adversarial networks for", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1806 }, { "text": "multi-domain image-to-image translation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 8789\u20138797, 2018. Chugh, K., Gupta, P., Dhall, A., and Subramanian, R. Not made for each other-audio-visual dissonance-based deep- fake detection and localization. pp. 439\u2013447, 2020.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1807 }, { "text": "Cohn, M. and Zellou, G. Perception of Concatenative vs. Neural Text-To-Speech (TTS): Differences in Intelligibil- ity in Noise and Language Attitudes. In Proc. Interspeech 2020, pp. 1733\u20131737, 2020. doi: 10.21437/Interspeech. 2020-1336. Corvi, R., Cozzolino, D., Zingarini, G., Poggi, G., Nagano, K., and Verdoliva, L. On the detection of synthetic images", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1808 }, { "text": "generated by diffusion models. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 1\u20135. IEEE, 2023. Dale, K., Sunkavalli, K., Johnson, M. K., Vlasic, D., Matusik, W., and Pfister, H. Video face replacement. Proceedings of the 2011 SIGGRAPH Asia Conference,", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1809 }, { "text": "2011. URL https://api.semanticscholar. org/CorpusID:8692593. Dolhansky, B., Bitton, J., Pflaum, B., Lu, J., Howes, R., Wang, M., and Ferrer, C. C. The deepfake detection challenge (dfdc) dataset, 2020. Donahue, C., McAuley, J., and Puckette, M. Adversarial audio synthesis, 2019. Fan, B., Wang, L., Soong, F. K., and Xie, L. Photo-real", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1810 }, { "text": "talking head with deep bidirectional lstm. In 2015 IEEE International Conference on Acoustics, Speech and Sig- nal Processing (ICASSP), pp. 4884\u20134888. IEEE, 2015. Fan, Y., Lin, Z., Saito, J., Wang, W., and Komura, T. Face- former: Speech-driven 3d facial animation with trans- formers. In Proceedings of the IEEE/CVF Conference", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1811 }, { "text": "on Computer Vision and Pattern Recognition, pp. 18770\u2013 18780, 2022. Frank, J. and Sch\u00a8onherr, L. Wavefake: A data set to facilitate audio deepfake detection, 2021. Garrido, P., Valgaerts, L., Rehmsen, O., Thorm\u00a8ahlen, T., P\u00b4erez, P., and Theobalt, C. Automatic face reenactment. CoRR, abs/1602.02651, 2016. URL http://arxiv.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1812 }, { "text": "org/abs/1602.02651. Ghosh, S., Seth, A., Umesh, S., and Manocha, D. Mast: Multiscale audio spectrogram transformers, 2023. Gimeno, P., Vi\u02dcnals, I., Ortega, A., Miguel, A., and Lleida, E. Multiclass audio segmentation based on recurrent neural networks for broadcast domain data. EURASIP Journal on Audio, Speech, and Music Processing, 2020", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1813 }, { "text": "(1), March 2020. ISSN 1687-4722. doi: 10.1186/ s13636-020-00172-6. URL http://dx.doi.org/ 10.1186/s13636-020-00172-6. Gong, Y., Chung, Y.-A., and Glass, J. Ast: Audio spectro- gram transformer, 2021. Guera, David; Delp, E. J. Deepfake video detection using recurrent neural networks. IEEE 2018 15th IEEE Interna-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1814 }, { "text": "tional Conference on Advanced Video and Signal Based Surveillance (AVSS), 2018. Guo, H., Hu, S., Wang, X., Chang, M.-C., and Lyu, S. Eyes tell all: Irregular pupil shapes reveal gan-generated faces. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1815 }, { "text": "2904\u20132908. IEEE, 2022. Guo, Y., Zhang, L., Hu, Y., He, X., and Gao, J. Ms-celeb-1m: A dataset and benchmark for large-scale face recognition. In Computer Vision\u2013ECCV 2016: 14th European Confer- ence, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part III 14, pp. 87\u2013102. Springer, 2016.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1816 }, { "text": "Guo, Y., Chen, K., Liang, S., Liu, Y., Bao, H., and Zhang, J. Ad-nerf: Audio driven neural radiance fields for talking head synthesis. In IEEE/CVF International Conference on Computer Vision (ICCV), 2021. Hamza, A., Javed, A. R. R., Iqbal, F., Kryvinska, N., Al- madhor, A. S., Jalil, Z., and Borghol, R. Deepfake audio", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1817 }, { "text": "detection via mfcc features using machine learning. IEEE Access, 10:134018\u2013134028, 2022. ISSN 2169-3536. doi: 10.1109/access.2022.3231480. URL http://dx.doi. org/10.1109/ACCESS.2022.3231480. Hande, S. S. A review of concatenative text to speech synthesis. International Journal of Latest Trends in Engineering and Technology (IJLTEMAS), 3(9):12\u2013", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1818 }, { "text": "The Tug-of-War Between Deepfake Generation and Detection He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learn- ing for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770\u2013778, 2016. Ho, J., Jain, A., and Abbeel, P. Denoising diffu- sion probabilistic models.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1819 }, { "text": "ArXiv, abs/2006.11239, 2020. URL https://api.semanticscholar. org/CorpusID:219955663. Ho, J., Chan, W., Saharia, C., Whang, J., Gao, R., Gritsenko, A., Kingma, D. P., Poole, B., Norouzi, M., Fleet, D. J., et al. Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303,", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1820 }, { "text": "2022. Hong, W., Ding, M., Zheng, W., Liu, X., and Tang, J. Cogvideo: Large-scale pretraining for text-to-video gener- ation via transformers. arXiv preprint arXiv:2205.15868, 2022. Hou, Y., Guo, Q., Huang, Y., Xie, X., Ma, L., and Zhao, J. Evading deepfake detectors via adversarial statistical consistency. In Proceedings of the IEEE/CVF Conference", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1821 }, { "text": "on Computer Vision and Pattern Recognition, pp. 12271\u2013 12280, 2023. Hsu, T. Fake and explicit images of taylor swift started on 4chan, study says, 2024. Hu, C., Tian, Q., Li, T., Yuping, W., Wang, Y., and Zhao, H. Neural dubber: Dubbing for videos ac- cording to scripts. In Ranzato, M., Beygelzimer, A.,", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1822 }, { "text": "Dauphin, Y., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems, volume 34, pp. 16582\u201316595. Curran Associates, Inc., 2021. URL https://proceedings.neurips. cc/paper_files/paper/2021/file/ 8a9c8ac001d3ef9e4ce39b1177295e03-Paper. pdf. Huang, G. B., Ramesh, M., Berg, T., and Learned-Miller,", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1823 }, { "text": "E. Labeled faces in the wild: A database for studying face recognition in unconstrained environments. Techni- cal Report 07-49, University of Massachusetts, Amherst, October 2007. Huzaifah, M. and Wyse, L. Deep generative models for musical audio synthesis, 2020. Jamaludin, A., Chung, J. S., and Zisserman, A. You said", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1824 }, { "text": "that?: Synthesising talking faces from audio. Interna- tional Journal of Computer Vision, 127(11):1767\u20131779, 2019. Ji, X., Zhou, H., Wang, K., Wu, W., Loy, C. C., Cao, X., and Xu, F. Audio-driven emotional video portraits. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 14080\u201314089, 2021.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1825 }, { "text": "Jiang, L., Li, R., Wu, W., Qian, C., and Loy, C. C. Deeperforensics-1.0: A large-scale dataset for real-world face forgery detection, 2020. Kain, A. and Macon, M. W. Spectral voice conversion for text-to-speech synthesis. Proceedings of the 1998 IEEE International Conference on Acoustics, Speech and", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1826 }, { "text": "Signal Processing, ICASSP \u201998 (Cat. No.98CH36181), 1:285\u2013288 vol.1, 1998. URL https://api. semanticscholar.org/CorpusID:1332456. Karras, T., Aila, T., Laine, S., Herva, A., and Lehtinen, J. Audio-driven facial animation by joint end-to-end learning of pose and emotion. ACM Transactions on Graphics (TOG), 36(4):1\u201312, 2017a.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1827 }, { "text": "Karras, T., Aila, T., Laine, S., and Lehtinen, J. Progres- sive growing of gans for improved quality, stability, and variation. arXiv preprint arXiv:1710.10196, 2017b. Karras, T., Laine, S., and Aila, T. A style-based genera- tor architecture for generative adversarial networks. In Proceedings of the IEEE/CVF Conference on Computer", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1828 }, { "text": "Vision and Pattern Recognition, pp. 4401\u20134410, 2019. Karras, T., Laine, S., Aittala, M., Hellsten, J., Lehtinen, J., and Aila, T. Analyzing and improving the image quality of StyleGAN. In Proc. CVPR, 2020. Karras, T., Aittala, M., Laine, S., H\u00a8ark\u00a8onen, E., Hellsten, J., Lehtinen, J., and Aila, T. Alias-free generative adversarial", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1829 }, { "text": "networks. In Proc. NeurIPS, 2021. Kawar, B., Zada, S., Lang, O., Tov, O., Chang, H.- T., Dekel, T., Mosseri, I., and Irani, M. Imagic: Text-based real image editing with diffusion mod- els. 2023 IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition (CVPR), pp. 6007\u20136017, 2022. URL https://api.semanticscholar.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1830 }, { "text": "org/CorpusID:252918469. Korshunov, P. and Marcel, S. Deepfakes: A new threat to face recognition? assessment and detection. arXiv preprint arXiv:1812.08685, 2018. Li, L., Bao, J., Zhang, T., Yang, H., Chen, D., Wen, F., and Guo, B. Face x-ray for more general face forgery detection. In Proceedings of the IEEE/CVF conference on", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1831 }, { "text": "computer vision and pattern recognition, pp. 5001\u20135010, 2020a. Li, Y. and Lyu, S. Exposing deepfake videos by detecting face warping artifacts. arXiv preprint arXiv:1811.00656, 2018. Li, Y., Chang, M.-C., and Lyu, S. In ictu oculi: Exposing ai created fake videos by detecting eye blinking. pp. 1\u20137,", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1832 }, { "text": "The Tug-of-War Between Deepfake Generation and Detection Li, Y., Yang, X., Sun, P., Qi, H., and Lyu, S. Celeb-df: A large-scale challenging dataset for deepfake forensics, 2020b. Liang, B., Pan, Y., Guo, Z., Zhou, H., Hong, Z., Han, X., Han, J., Liu, J., Ding, E., and Wang, J. Expressive talking head generation with granular audio-visual control. In", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1833 }, { "text": "Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3387\u20133396, 2022. Lim, Y., Lee, C., Kim, A., and Etzioni, O. Distildire: A small, fast, cheap and lightweight diffusion synthesized deepfake detection, 2024. URL https://arxiv. org/abs/2406.00856. Lin, Y., Wang, S., Lin, Q., and Tang, F.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1834 }, { "text": "Face swap- ping under large pose variations: A 3d model based ap- proach. 2012 IEEE International Conference on Multime- dia and Expo, pp. 333\u2013338, 2012. URL https://api. semanticscholar.org/CorpusID:30759641. Lin, Y., Shen, J., Wang, Y., and Pantic, M. Fp-age: Leverag- ing face parsing attention for facial age estimation in the", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1835 }, { "text": "wild. arXiv, 2021. Liu, S., Cao, Y., and Meng, H. M. Emotional voice conversion with cycle-consistent adversarial network. arXiv: Audio and Speech Processing, 2020a. URL https://api.semanticscholar. org/CorpusID:215416116. Liu, X., Xu, Y., Wu, Q., Zhou, H., Wu, W., and Zhou, B. Semantic-aware implicit neural audio-driven video", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1836 }, { "text": "portrait generation. ECCV, 2022. Liu, X., Lu, H., Yuan, J., and Li, X. Cat: Causal audio transformer for audio classification, 2023a. Liu, Z., Luo, P., Wang, X., and Tang, X. Deep learning face attributes in the wild. In Proceedings of International Conference on Computer Vision (ICCV), December 2015.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1837 }, { "text": "Liu, Z., Qi, X., and Torr, P. H. Global texture enhancement for fake face detection in the wild. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 8060\u20138069, 2020b. Liu, Z., Li, M., Zhang, Y., Wang, C., Zhang, Q., Wang, J., and Nie, Y. Fine-grained face swapping via regional gan", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1838 }, { "text": "inversion, 2023b. Lorenz, P., Durall, R. L., and Keuper, J. Detecting images generated by deep diffusion models using their local in- trinsic dimensionality. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 448\u2013 459, 2023. Lu, S.-A. shaoanlu/faceswap-GAN, May 2024. URL", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1839 }, { "text": "https://github.com/shaoanlu/ faceswap-GAN. original-date: 2017-12- 23T08:40:32Z. Lu, Y. and Ebrahimi, T. Towards the detection of ai-synthesized human face images. arXiv preprint arXiv:2402.08750, 2024. Lu, Y., Chai, J., and Cao, X. Live speech portraits: real-time photorealistic talking-head animation. ACM Transactions", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1840 }, { "text": "on Graphics (TOG), 40(6):1\u201317, 2021. Lv, L. Smart watermark to defend against deepfake image manipulation. In 2021 IEEE 6th international conference on computer and communication systems (ICCCS), pp. 380\u2013384. IEEE, 2021. Marek. MarekKowalski/FaceSwap, May 2024. URL https://github.com/MarekKowalski/", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1841 }, { "text": "FaceSwap. original-date: 2016-06-19T00:09:07Z. Marra, F., Gragnaniello, D., Verdoliva, L., and Poggi, G. Do gans leave artificial fingerprints? In 2019 IEEE confer- ence on multimedia information processing and retrieval (MIPR), pp. 506\u2013511. IEEE, 2019. Matern, F., Riess, C., and Stamminger, M. Exploiting vi-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1842 }, { "text": "sual artifacts to expose deepfakes and face manipulations. In 2019 IEEE Winter Applications of Computer Vision Workshops (WACVW), pp. 83\u201392. IEEE, 2019. Mirsky, Y. and Lee, W. The creation and detection of deep- fakes: A survey. ACM Computing Surveys, 54(1):1\u201341, January 2021. ISSN 1557-7341. doi: 10.1145/3425780.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1843 }, { "text": "URL http://dx.doi.org/10.1145/3425780. M\u00a8uller, N. M., Czempin, P., Dieckmann, F., Froghyar, A., and B\u00a8ottinger, K. Does audio deepfake detection general- ize?, 2022. Neekhara, P., Dolhansky, B., Bitton, J., and Ferrer, C. C. Adversarial threats to deepfake detection: A practical perspective. In Proceedings of the IEEE/CVF conference", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1844 }, { "text": "on computer vision and pattern recognition, pp. 923\u2013932, 2021. Nguyen, T. T., Nguyen, Q. V. H., Nguyen, D. T., Nguyen, D. T., Huynh-The, T., Nahavandi, S., Nguyen, T. T., Pham, Q.-V., and Nguyen, C. M. Deep learning for deepfakes creation and detection: A survey. Computer Vision and Image Understanding, 223:103525, Octo-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1845 }, { "text": "The Tug-of-War Between Deepfake Generation and Detection abs/1908.05932, 2019. URL http://arxiv.org/ abs/1908.05932. Nirkin, Y., Keller, Y., and Hassner, T. Fsganv2: Improved subject agnostic face swapping and reenactment. 2022. Ojha, U., Li, Y., and Lee, Y. J. Towards universal fake image detectors that generalize across generative models. In", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1846 }, { "text": "Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 24480\u201324489, 2023. Park, S. J., Kim, M., Hong, J., Choi, J., and Ro, Y. M. Synctalkface: Talking face generation with precise lip- syncing via audio-lip memory. In AAAI Conference on Artificial Intelligence. Association for the Advancement", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1847 }, { "text": "of Artificial Intelligence, 2022. Perov, I., Gao, D., Chervoniy, N., Liu, K., Marangonda, S., Um\u00b4e, C., Dpfks, M., Facenheim, C. S., RP, L., Jiang, J., Zhang, S., Wu, P., Zhou, B., and Zhang, W. Deep- facelab: A simple, flexible and extensible face swap- ping framework. CoRR, abs/2005.05535, 2020. URL", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1848 }, { "text": "https://arxiv.org/abs/2005.05535. Pianese, A., Cozzolino, D., Poggi, G., and Verdoliva, L. Training-free deepfake voice recognition by leveraging large-scale pre-trained models, 2024. Prajwal, K., Mukhopadhyay, R., Namboodiri, V. P., and Jawahar, C. A lip sync expert is all you need for speech to lip generation in the wild. In Proceedings of the 28th", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1849 }, { "text": "ACM International Conference on Multimedia, pp. 484\u2013 492, 2020. Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pp. 8748\u20138763. PMLR, 2021.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1850 }, { "text": "Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I. Zero- shot text-to-image generation. ArXiv, abs/2102.12092, 2021. URL https://api.semanticscholar. org/CorpusID:232035663. Rana, M. S., Nobi, M. N., Murali, B., and Sung, A. H. Deepfake detection: A systematic literature review. IEEE", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1851 }, { "text": "Access, 10:25494\u201325513, 2022. doi: 10.1109/ACCESS. 2022.3154404. Rashid, M. M., Lee, S.-H., and Kwon, K.-R. Blockchain technology for combating deepfake and protect video/image integrity. Journal of Korea Multimedia Soci- ety, 24(8):1044\u20131058, 2021. Reimao, R. and Tzerpos, V. For: A dataset for syn-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1852 }, { "text": "thetic speech detection. In 2019 International Confer- ence on Speech Technology and Human-Computer Dia- logue (SpeD), pp. 1\u201310, 2019. doi: 10.1109/SPED.2019. 8906599. Reisach, U. The responsibility of social media in times of societal and political manipulation. European journal of operational research, 291(3):906\u2013917, 2021.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1853 }, { "text": "Richard, A., Zollh\u00a8ofer, M., Wen, Y., de la Torre, F., and Sheikh, Y. Meshtalk: 3d face animation from speech using cross-modality disentanglement. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021. Richardson, E., Alaluf, Y., Patashnik, O., Nitzan, Y., Azar, Y., Shapiro, S., and Cohen-Or, D. Encoding in style:", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1854 }, { "text": "a stylegan encoder for image-to-image translation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2287\u20132296, 2021. Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. 2022 IEEE/CVF Confer-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1855 }, { "text": "ence on Computer Vision and Pattern Recognition (CVPR), pp. 10674\u201310685, 2021. URL https: //api.semanticscholar.org/CorpusID: 245335280. R\u00a8ossler, A., Cozzolino, D., Verdoliva, L., Riess, C., Thies, J., and Nie\u00dfner, M. FaceForensics++: Learning to detect manipulated facial images. In International Conference", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1856 }, { "text": "on Computer Vision (ICCV), 2019. Rothe, R., Timofte, R., and Gool, L. V. Deep expectation of real and apparent age from a single image without facial landmarks. International Journal of Computer Vision, 126(2-4):144\u2013157, 2018. Ruhs, H. Facefusion, May 2024. URL https: //github.com/facefusion/facefusion.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1857 }, { "text": "original-date: 2023-08-17T19:59:55Z. Ruiz, N., Li, Y., Jampani, V., Pritch, Y., Rubinstein, M., and Aberman, K. Dreambooth: Fine tuning text- to-image diffusion models for subject-driven genera- tion. 2023 IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition (CVPR), pp. 22500\u201322510, 2022.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1858 }, { "text": "The Tug-of-War Between Deepfake Generation and Detection Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E. L., Ghasemipour, S. K. S., Ayan, B. K., Mahdavi, S. S., Lopes, R. G., Salimans, T., Ho, J., Fleet, D. J., and Norouzi, M. Pho- torealistic text-to-image diffusion models with deep", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1859 }, { "text": "language understanding. ArXiv, abs/2205.11487, 2022. URL https://api.semanticscholar. org/CorpusID:248986576. Scarpiniti, M., Comminiello, D., Uncini, A., and Lee, Y.-C. Deep recurrent neural networks for audio classification in construction sites. In 2020 28th European Signal Pro- cessing Conference (EUSIPCO), pp. 810\u2013814, 2021. doi:", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1860 }, { "text": "10.23919/Eusipco47968.2020.9287802. Shen, J., Pang, R., Weiss, R. J., Schuster, M., Jaitly, N., Yang, Z., Chen, Z., Zhang, Y., Wang, Y., Skerry- Ryan, R. J., Saurous, R. A., Agiomyrgiannakis, Y., and Wu, Y. Natural tts synthesis by condition- ing wavenet on mel spectrogram predictions. 2018 IEEE International Conference on Acoustics, Speech", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1861 }, { "text": "and Signal Processing (ICASSP), pp. 4779\u20134783, 2017. URL https://api.semanticscholar. org/CorpusID:206742911. Shen, S., Li, W., Zhu, Z., Duan, Y., Zhou, J., and Lu, J. Learning dynamic facial radiance fields for few-shot talk- ing head synthesis. In ECCV, 2022. Sisman, B., Yamagishi, J., King, S., and Li, H. An overview", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1862 }, { "text": "of voice conversion and its challenges: From statistical modeling to deep learning. August 2020. Stypu\u0142kowski, M., Vougioukas, K., He, S., Zieba, M., Petridis, S., and Pantic, M. Diffused heads: Diffusion models beat gans on talking-face generation. pp. 5091\u2013 5100, 2024. Sun, Y., Zhang, Z., Echizen, I., Nguyen, H. H., Qiu, C.,", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1863 }, { "text": "and Sun, L. Face forgery detection based on facial re- gion displacement trajectory series. In Proceedings of the IEEE/CVF Winter Conference on Applications of Com- puter Vision, pp. 633\u2013642, 2023. Suwajanakorn, S., Seitz, S. M., and Kemelmacher- Shlizerman, I. Synthesizing obama: learning lip sync", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1864 }, { "text": "from audio. ACM Transactions on Graphics (ToG), 36 (4):1\u201313, 2017. Swathi, P. and Sk, S. Deepfake creation and detection: A survey. In 2021 Third International Conference on Inventive Research in Computing Applications (ICIRCA), pp. 584\u2013588. IEEE, 2021. Tan, X., Qin, T., Soong, F., and Liu, T.-Y. A survey on", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1865 }, { "text": "neural speech synthesis, 2021. Tariang, D., Corvi, R., Cozzolino, D., Poggi, G., Nagano, K., and Verdoliva, L. Synthetic image verification in the era of generative ai: What works and what isn\u2019t there yet. arXiv preprint arXiv:2405.00196, 2024. Thies, J., Zollh\u00a8ofer, M., Nie\u00dfner, M., Valgaerts, L., Stam-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1866 }, { "text": "minger, M., and Theobalt, C. Real-time expression trans- fer for facial reenactment. ACM Transactions on Graph- ics (TOG), 34(6), 2015. Thies, J., Zollhofer, M., Stamminger, M., Theobalt, C., and Nie\u00dfner, M. Face2face: Real-time face capture and reenactment of rgb videos. In Proceedings of the IEEE", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1867 }, { "text": "conference on computer vision and pattern recognition, pp. 2387\u20132395, 2016. Thies, J., Zollh\u00a8ofer, M., and Nie\u00dfner, M. Deferred neural rendering: Image synthesis using neural textures. ACM Transactions on Graphics (TOG), 38(4):1\u201312, 2019. Thies, J., Elgharib, M., Tewari, A., Theobalt, C., and Nie\u00dfner, M. Neural voice puppetry: Audio-driven fa-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1868 }, { "text": "cial reenactment. In European Conference on Computer Vision, pp. 716\u2013731. Springer, 2020. Tokuda, K., Yoshimura, T., Masuko, T., Kobayashi, T., and Kitamura, T. Speech parameter generation algorithms for hmm-based speech synthesis. 2000 IEEE International Conference on Acoustics, Speech, and Signal Processing.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1869 }, { "text": "Proceedings (Cat. No.00CH37100), 3:1315\u20131318 vol.3, 2000. URL https://api.semanticscholar. org/CorpusID:9372431. Tov, O., Alaluf, Y., Nitzan, Y., Patashnik, O., and Cohen-Or, D. Designing an encoder for stylegan image manipulation. ACM Transactions on Graphics (TOG), 2021. Trinh Van, L., Dao, T., Le, T., and Castelli, E. Emotional", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1870 }, { "text": "speech recognition using deep neural networks. Sensors, 22:1414, 02 2022. doi: 10.3390/s22041414. Umur Aybars Ciftci, Ilke Demir, L. Y. How do the hearts of deep fakes beat? deep fake source detection via interpret- ing residuals with biological signals. 2020. van den Oord, A., Dieleman, S., Zen, H., Simonyan,", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1871 }, { "text": "K., Vinyals, O., Graves, A., Kalchbrenner, N., Se- nior, A. W., and Kavukcuoglu, K. Wavenet: A gen- erative model for raw audio. ArXiv, abs/1609.03499, 2016. URL https://api.semanticscholar. org/CorpusID:6254678. Vlasic, D., Brand, M., Pfister, H., and Popovi\u00b4c, J. Face transfer with multilinear models. ACM Trans. Graph.,", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1872 }, { "text": "The Tug-of-War Between Deepfake Generation and Detection Wang, K., Wu, Q., Song, L., Yang, Z., Wu, W., Qian, C., He, R., Qiao, Y., and Loy, C. C. Mead: A large-scale audio- visual dataset for emotional talking-face generation. In ECCV, August 2020a. Wang, R., Juefei-Xu, F., Ma, L., Xie, X., Huang, Y., Wang,", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1873 }, { "text": "J., and Liu, Y. Fakespotter: A simple yet robust baseline for spotting ai-synthesized fake faces. arXiv preprint arXiv:1909.06122, 2019. Wang, S.-Y., Wang, O., Zhang, R., Owens, A., and Efros, A. A. Cnn-generated images are surprisingly easy to spot...for now. In CVPR, 2020b. Wang, X., Yamagishi, J., Todisco, M., Delgado, H., Nautsch,", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1874 }, { "text": "A., Evans, N., Sahidullah, M., Vestman, V., Kinnunen, T., Lee, K. A., Juvela, L., Alku, P., Peng, Y.-H., Hwang, H.-T., Tsao, Y., Wang, H.-M., Maguer, S. L., Becker, M., Henderson, F., Clark, R., Zhang, Y., Wang, Q., Jia, Y., Onuma, K., Mushika, K., Kaneda, T., Jiang, Y., Liu, L.-J., Wu, Y.-C., Huang, W.-C., Toda, T., Tanaka, K.,", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1875 }, { "text": "Kameoka, H., Steiner, I., Matrouf, D., Bonastre, J.-F., Govender, A., Ronanki, S., Zhang, J.-X., and Ling, Z.- H. Asvspoof 2019: A large-scale public database of synthesized, converted and replayed speech, 2020c. Wang, X., Li, Y., Zhang, H., and Shan, Y. Towards real- world blind face restoration with generative facial prior.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1876 }, { "text": "In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021. Wang, Z., Bao, J., Zhou, W., Wang, W., Hu, H., Chen, H., and Li, H. Dire for diffusion-generated image detection. In Proceedings of the IEEE/CVF International Confer- ence on Computer Vision, pp. 22445\u201322455, 2023. Wang, Z. J., Montoya, E., Munechika, D., Yang, H.,", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1877 }, { "text": "Hoover, B., and Chau, D. H. Diffusiondb: A large- scale prompt gallery dataset for text-to-image gener- ative models. arXiv:2210.14896 [cs], 2022. URL https://arxiv.org/abs/2210.14896. Wodajo, D. and Atnafu, S. Deepfake video detection using generative convolutional vision transformer. 2023. Wu, H.,", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1878 }, { "text": "Kuo, H.-C., Zheng, N., Hung, K.-H., yi Lee, H., Tsao, Y., Wang, H.-M., and Meng, H. M. Partially fake audio detection by self-attention- based fake span discovery. ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 9236\u20139240, 2022. URL https://api.semanticscholar.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1879 }, { "text": "org/CorpusID:246823684. Xu, C., Zhang, J., Hua, M., He, Q., Yi, Z., and Liu, Y. Region-aware face swapping. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7632\u20137641, 2022a. Xu, Y., Deng, B., Wang, J., Jing, Y., Pan, J., and He, S. High-resolution face swapping via latent semantics disen-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1880 }, { "text": "tanglement, 2022b. Xu, Z., Zhou, H., Hong, Z., Liu, Z., Liu, J., Guo, Z., Han, J., Liu, J., Ding, E., and Wang, J. Styleswap: Style-based generator empowers robust face swapping. In European Conference on Computer Vision, pp. 661\u2013677. Springer, 2022c. Yamagishi, J., Wang, X., Todisco, M., Sahidullah, M.,", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1881 }, { "text": "Patino, J., Nautsch, A., Liu, X., Lee, K. A., Kinnunen, T., Evans, N., and Delgado, H. Asvspoof 2021: accelerating progress in spoofed and deepfake speech detection, 2021. Yan, Z., Zhang, Y., Yuan, X., Lyu, S., and Wu, B. Deep- fakebench: A comprehensive benchmark of deepfake de- tection. In Thirty-seventh Conference on Neural Informa-", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1882 }, { "text": "tion Processing Systems Datasets and Benchmarks Track, 2023. URL https://openreview.net/forum? id=hizSx8pf0U. Yang, T., Ren, P., Xie, X., and Zhang, L. Gan prior em- bedded network for blind face restoration in the wild. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 672\u2013681, 2021.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1883 }, { "text": "Yang, X., Li, Y., and Lyu, S. Exposing deep fakes using inconsistent head poses. In ICASSP 2019-2019 IEEE In- ternational Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 8261\u20138265. IEEE, 2019. Yang, Y., Shillingford, B., Assael, Y., Wang, M., Liu, W., Chen, Y., Zhang, Y., Sezener, E., Cobo, L. C., Denil,", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1884 }, { "text": "M., et al. Large-scale multilingual audio visual dubbing. arXiv preprint arXiv:2011.03530, 2020. Yi, J., Fu, R., Tao, J., Nie, S., Ma, H., Wang, C., Wang, T., Tian, Z., Bai, Y., Fan, C., Liang, S., Wang, S., Zhang, S., Yan, X., Xu, L., Wen, Z., Li, H., Lian, Z., and Liu, B. Add 2022: the first audio deep synthesis detection", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1885 }, { "text": "challenge, 2022. Yi, J., Tao, J., Fu, R., Yan, X., Wang, C., Wang, T., Zhang, C. Y., Zhang, X., Zhao, Y., Ren, Y., Xu, L., Zhou, J., Gu, H., Wen, Z., Liang, S., Lian, Z., Nie, S., and Li, H. Add 2023: the second audio deepfake detection challenge, 2023a. Yi, J., Wang, C., Tao, J., Zhang, X., Zhang, C. Y., and Zhao,", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1886 }, { "text": "Y. Audio deepfake detection: A survey, 2023b. Yi, J., Wang, C., Tao, J., Zhang, C. Y., Fan, C., Tian, Z., Ma, H., and Fu, R. Scenefake: An initial dataset and benchmarks for scene fake audio detection, 2024. Yu, F., Seff, A., Zhang, Y., Song, S., Funkhouser, T., and Xiao, J. Lsun: Construction of a large-scale image dataset", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1887 }, { "text": "The Tug-of-War Between Deepfake Generation and Detection using deep learning with humans in the loop. arXiv preprint arXiv:1506.03365, 2015. Yu, P., Xia, Z., Fei, J., and Lu, Y. A survey on deepfake video detection. IET Biometrics, 10(6):607\u2013624, April 2021. ISSN 2047-4946. doi: 10.1049/bme2.12031. URL", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1888 }, { "text": "http://dx.doi.org/10.1049/bme2.12031. Zaman, K., Sah, M., Direkoglu, C., and Unoki, M. A survey of audio classification using deep learning. IEEE Access, 11:106620\u2013106649, 2023. doi: 10.1109/ACCESS.2023. 3318015. Zhang, X., Karaman, S., and Chang, S.-F. Detecting and simulating artifacts in gan fake images. In 2019 IEEE", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1889 }, { "text": "international workshop on information forensics and se- curity (WIFS), pp. 1\u20136. IEEE, 2019. Zhang, Z., Hu, Z., Deng, W., Fan, C., Lv, T., and Ding, Y. Dinet: Deformation inpainting network for realistic face visually dubbing on high resolution video. AAAI, 2023. Zhao, Y., Yi, J., Tao, J., Wang, C., Zhang, X., and Dong,", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1890 }, { "text": "Y. Emofake: An initial dataset for emotion fake audio detection, 2023. Zhou, X. and Zafarani, R. A survey of fake news: Fun- damental theories, detection methods, and opportuni- ties, September 2020. ISSN 1557-7341. URL http: //dx.doi.org/10.1145/3395046. Zhou, Y., Xu, Z., Landreth, C., Kalogerakis, E., Maji, S.,", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1891 }, { "text": "and Singh, K. Visemenet: Audio-driven animator-centric speech animation. ACM Transactions on Graphics (TOG), 37(4):1\u201310, 2018. Zhou, Y., Han, X., Shechtman, E., Echevarria, J., Kaloger- akis, E., and Li, D. Makelttalk: speaker-aware talking- head animation. ACM Transactions on Graphics (TOG), 39(6):1\u201315, 2020.", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1892 }, { "text": "Zhu, J.-Y., Park, T., Isola, P., and Efros, A. A. Unpaired image-to-image translation using cycle-consistent adver- sarial networks. In Proceedings of the IEEE international conference on computer vision, pp. 2223\u20132232, 2017. Zi, B., Chang, M., Chen, J., Ma, X., and Jiang, Y.-G. Wild- deepfake: A challenging real-world dataset for deepfake", "source": "Tug of War Deepfake Detection", "year": 2024, "url": "https://arxiv.org/abs/2407.06174", "id": 1893 }, { "text": "IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. 50, NO. 1, NOVEMBER 2024 1 Deepfake Media Generation and Detection in the Generative AI Era: A Survey and Outlook Florinel-Alin Croitoru, Andrei-Iulian H\u02c6\u0131ji, Vlad Hondru, Nicolae C\u02d8at\u02d8alin Ristea, Paul Irofti, Marius Popescu, Cristian Rusu, Radu Tudor Ionescu, Fahad Shahbaz Khan, Senior Member, IEEE,", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1894 }, { "text": "and Mubarak Shah, Fellow, IEEE Abstract\u2014With the recent advancements in generative modeling, the realism of deepfake content has been increasing at a steady pace, even reaching the point where people often fail to detect manipulated media content online, thus being deceived into various kinds of scams. In this paper, we survey deepfake generation and detection techniques, including the most recent developments in the", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1895 }, { "text": "field, such as diffusion models and Neural Radiance Fields. Our literature review covers all deepfake media types, comprising image, video, audio and multimodal (audio-visual) content. We identify various kinds of deepfakes, according to the procedure used to alter or generate the fake content. We further construct a taxonomy of deepfake generation and detection methods, illustrating the important", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1896 }, { "text": "groups of methods and the domains where these methods are applied. Next, we gather datasets used for deepfake detection and provide updated rankings of the best performing deepfake detectors on the most popular datasets. In addition, we develop a novel multimodal benchmark to evaluate deepfake detectors on out-of-distribution content. The results indicate that state-of-the-art detectors", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1897 }, { "text": "fail to generalize to deepfake content generated by unseen deepfake generators. Finally, we propose future directions to obtain robust and powerful deepfake detectors. Our project page and new benchmark are available at https://github.com/CroitoruAlin/biodeep. Index Terms\u2014deepfake, deepfake generation, deepfake detection, deepfake benchmark.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1898 }, { "text": "\u2726 1 INTRODUCTION D EEPFAKE media comprises image, video or audio files that are digitally altered or generated from scratch with AI tools in order to impersonate real or non-existent people. The recent groundbreaking progress of generative AI methods [1]\u2013[6] has enabled the creation of realistic deep-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1899 }, { "text": "fake media with very little effort [7]\u2013[18]. Unfortunately, the generated deepfake media can be used by scammers to spread misinformation on social media platforms to achieve large-scale political manipulation, and to deceive individu- als or companies into financial frauds. In an age where misinformation can quickly spread", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1900 }, { "text": "through social media platforms, deepfakes pose a critical threat to public trust and democracy, especially due to their growing online exploitation. A recent analysis of the fraud trends indicates that the number of fraud cases based on deepfakes registered a 10\u00d7 increase in 2023, with respect to \u2022", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1901 }, { "text": "F.A. Croitoru, A.I. H\u02c6\u0131ji, V. Hondru, N.C. Ristea, P. Irofti, M. Popescu, C. Rusu and R.T. Ionescu are with the Department of Computer Science, University of Bucharest, Bucharest, Romania. F.A. Croitoru, A.I. H\u02c6\u0131ji, V. Hondru, and N.C. Ristea have contributed equally. R.T. Ionescu is the corresponding author.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1902 }, { "text": "E-mail: raducu.ionescu@gmail.com \u2022 F.S. Khan is with Mohamed bin Zayed University of Artificial Intelligence (MBZUAI), UAE, and Link\u00a8oping University, Sweden. \u2022 M. Shah is with the Center for Research in Computer Vision (CRCV), Department of Computer Science, University of Central Florida, Orlando,", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1903 }, { "text": "FL, US. Copyright 2024 IEEE. Personal use of this material is permitted. Permis- sion from IEEE must be obtained for all other uses, including reprint- ing/republishing this material for advertising or promotional purposes, collect- ing new collected works for resale or redistribution to servers or lists, or reuse", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1904 }, { "text": "of any copyrighted component of this work in other works. This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible. 20221. Another recent study found that about 70% of people are unable to distinguish between a real and a deepfake", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1905 }, { "text": "voice2. The growing quality and quantity of deepfakes raise significant concerns, particularly regarding online fraud and manipulation. To prevent the spread of deepfake media, researchers have developed a broad range of unimodal [19]\u2013[23] or multimodal [24]\u2013[26] methods for deepfake detection. However, deepfake detectors trained on media", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1906 }, { "text": "generated with a certain set of AI tools typically fail on deepfakes generated with a distinct set of tools [20]\u2013[22]. This has led to a relentless race to develop more powerful and robust deepfake detectors. To this end, we conduct a comprehensive survey on the recent developments in deepfake media generation and", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1907 }, { "text": "detection. We first define a set of deepfake categories, which are determined based on the procedure used to generate the deepfake content. We identify both domain-agnostic and domain-specific deepfake types, and explain what kind of deepfake media belongs to each category. We next build a taxonomy of deepfake generation and detection methods,", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1908 }, { "text": "which creates a multi-perspective hierarchical categoriza- tion based on the considered media types, the employed architectures and the targeted tasks. As shown in Figure 1, we first divide contributions by task, into generation and detection. For each task, we identify the employed architectures. For deepfake generation, we find that the most", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1909 }, { "text": "popular architectures are Generative Adversarial Networks (GANs) [8], [14]\u2013[16], [27], [28] and denoising diffusion models [11]\u2013[13], [18], [29]\u2013[31]. To detect deepfakes, the 1. Sumsub Expert Roundtable: The Top KYC Trends Coming in 2024 2. Artificial Imposters\u2013Cybercriminals Turn to AI Voice Cloning for a", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1910 }, { "text": "IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. 50, NO. 1, NOVEMBER 2024 2 Deepfake Generation Detection GAN Diffusion Transformer NeRF VAE 3DMM Zhang et al. ICASSP 2022 Doukas et al. ICCV 2021 Wang et al. CVPR 2023b Blattmann et al. CVPR 2023 Ma et al. ArXiv 2024a Zhang et al. ArXiv 2024", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1911 }, { "text": "Bounareli et al. ArXiv 2024 Singer et al. ICLR 2022 Ho et al. ArXiv 2022a Yang et al. ArXiv 2024 Blattmann et al. ArXiv 2023 Ma et al. ArXiv 2024b Bao et al. ArXiv 2024 Guo et al. ICLR 2024 Wu et al. ICCV 2023 Wang et al. ArXiv 2024 Ho et al. ArXiv 2022b Huang et al. IJCAI 2023 Huang et al. ACMMM 2022", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1912 }, { "text": "Shen et al. ICLR 2024 Ju et al. ArXiv 2024 Du et al. AAAI 2024 Yang et al. TASLP 2024 Ren et al. ICLR 2021 Wei et al. ArXiv 2024 Chen et al. ArXiv 2024a Stypu\u0142kowski et al. WACV 2024 Xu et al. ArXiv 2024a Tian et al. ArXiv 2024 Xu et al. ArXiv 2024b Wang et al. ArXiv 2024a Rochow et al. CVPR 2024 Yan et al. ArXiv 2021", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1913 }, { "text": "Hong et al. ICLR 2023 Villegas et al. ICLR 2023 Yu et al. CVPR 2023 Jiang et al. ICCV 2023 Ren et al. NeurIPS 2019 Wang et al. ArXiv 2023 Jiang et al. ArXiv 2023 Kharitonov et al. TACL 2023 Yang et al. ArXiv 2023 Jang et al. CVPR 2024 Cheng et al. SIGGRAPH 2022 Wang et al. AAAI 2022 Ling et al. JSTSP 2023", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1914 }, { "text": "Gan et al. ICCV 2023 Zhang et al. CVPR 2023b Tan et al. TPAMI 2024 Li et al. CVPR 2023b Peng et al. CVPR 2024a Ye et al. ArXiv 2023 Yang et al. ECCV 2022 Thies et al. CVPR 2016 CNN Korshunova et al. ICCV 2017 Wang et al. IJCAI 2021 Thies et al. TOG 2019 Hwang et al. ICASSP 2023 Bao et al. ICCV 2017", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1915 }, { "text": "GAN+VAE Kim et al. ICML 2021 Casanova et al. ICML 2022 Lee et al. ICLR 2023 CNN+RNN Liu et al. NeurIPS 2022 Lu et al. TOG 2021 Gururani et al. ICCV 2023 Peng et al. ICCV 2023 Zhong et al. CVPR 2023 Haliassos et al. CVPR 2021 Haliassos et al. CVPR 2022 Agarwal et al. WIFS 2020 Zhao et al. ICICS 2020", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1916 }, { "text": "Wang et al. AAAI 2023 Gu et al. AAAI 2022 Gu et al. ECCV 2022 Cozzolino et al. ICCV 2021 Demir et al. WACV 2024 Bonettini et al. ICPR 2021 Yan et al. ArXiv 2024 Tak et al. ICASSP 2021 Wang et al. INTERSPEECH 2023 Conti et al. ICASSP 2022 Hua et al. SPL 2021 Zhang et al. AAAI 2024 Wang et al. ACMMM 2020", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1917 }, { "text": "Raza et al. CVPRW 2023 Cozzolino et al. CVPRW 2023 Kihal et al. MTA 2023 Dong et al. CVPR 2022 Wang et al. ICMR 2022 Aghasanli et al. ICCVW 2023 Cai et al. CVPR 2023 Xu et al. IJCV 2024 Bartusiak et al. ACSSC 2021 Bartusiak et al. ICMLA 2022 Liu et al. ICASSP 2023 Cai et al. ICASSP 2023 Zhang et al. IHMMSec 2021", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1918 }, { "text": "Martin et al. ICASSP 2022 Jung et al. ICASSP 2022 Oorloff et al. CVPR 2024 Salvi et al. JI 2023 Ilyas et al. ASC 2023 Asha et al. MS 2024 Liu et al. SPIC 2023 Yang et al. TIFS 2023b Feng et al. CVPR 2023 Zou et al. ICASSP 2024 Nie et al. ACMMM 2024 Zhou et al. ICCV 2021 Transformer RNN CNN+RNN Hu et al. AAAI 2022", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1919 }, { "text": "Masi et al. ECCV 2020 Sabir et al. CVPRW 2019 Montserrat et al. CVPRW 2020 G\u00fcera et al. AVSS 2018 Liu et al. WACV 2023 Amerini et al. ACM 2020 Goyal et al. TCSS 2023 CNN+Transformer Hong et al. CVPR 2024 Shao et al. ECCV 2022 Kamat et al. ICCVW 2023 Zheng et al. ICCV 2021 Choi et al. CVPR 2024 Xu et al. ICCV 2023", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1920 }, { "text": "Guan et al. NeurIPS 2022 Coccomini et al. ICIAP 2022 Wang et al. CVPR 2023 Yang et al. TIFS 2023a Cao et al. CVPR 2022 Tan et al. AAAI 2023 Jung et al. ICASSP 2022 Chen et al. ICASSP 2023 Tak et al. ASVSPOOF 2021 Tak et al. INTERSPEECH 2021 GNN Image Video Audio Multimodal Legend: CNN Face synthesis:", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1921 }, { "text": "Boutros et al. ICCV 2023 Chen et al. ICLR 2024 Kim et al. ArXiv 2024 Papantoniou et al. ECCV 2024 Podell et al. ICLR 2024 Wang et al. CVPR 2024 Huang et al. ArXiv 2024a Wang et al. ArXiv 2024b Wang et al. ArXiv 2024d Saharia et al. NeurIPS 2022 Ruiz et al. CVPR 2023 Rombach et al. CVPR 2022 Diffusion+GAN", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1922 }, { "text": "Li et al. CVPR 2024b Wang et al. ArXiv 2024e Nirkin et al. ICCV 2019 Xu et al. ECCV 2022 Gao et al. CVPR 2023b Skorokhodov et al. CVPR 2022 Tulyakov et al. CVPR 2018 Oorloff et al. ICCV 2023 Yu et al. ICLR 2022 Brooks et al. NeurIPS 2022 Tian et al. ICLR 2021 Traditional Machine Learning Lugstein et al. IHMMSec 2021", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1923 }, { "text": "Guarnera et al. CVPRW 2020 Hosler et al. CVPRW 2021 Face swapping: Zhao et al. CVPR 2023 Gu et al. ECCV 2024 Huang et al. ArXiv 2024b Han et al. ArXiv 2023 Attribute manipulation: Zhu et al. CVPR 2017 Bao et al. CVPR 2018 Choi et al. CVPR 2018 Tov et al. TOG 2021 Hsu et al. CVPR 2022 Suwa\u0142a et al. WACV 2024", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1924 }, { "text": "Bounareli et al. ICCV 2023 Attribute manipulation: Liu et al. ArXiv 2023 Chen et al. AAAI 2024 Chen et al. ArXiv 2024b Guo et al. NeurIPS 2024 He et al. ArXiv 2024 Lin et al. CVPR 2024a Liu et al. CVPR 2024 Ma et al. SIGGRAPH 2024 Peng et al. CVPR 2024b Wu et al. ArXiv 2024 Li et al. CVPR 2024a Wang et al. ArXiv 2024c", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1925 }, { "text": "Wei et al. ECCV 2024 Face synthesis: Karras et al. ICLR 2017 Shen et al. CVPR 2018 Karras et al. TPAMI 2021 Choi et al. CVPR 2020 Karras et al. CVPR 2020 Esser et al. CVPR 2021 Karras et al. NeurIPS 2021 Fu et al. ECCV 2022 Sauer et al. SIGGRAPH 2022 Xu et al. CVPR 2023 Face swapping: Moniz et al. NeurIPS 2018", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1926 }, { "text": "Sun et al. ECCV 2018 Chen et al. ACMMM 2020 Li et al. CVPR 2020 Chen et al. CVPR 2021 Gao et al. CVPR 2021 Zhu et al. CVPR 2021 Kim et al. CVPR 2022 Cao et al. FG 2023 Li et al. CVPR 2023a Liu et al. CVPR 2023 Ren et al. ICCV 2023 Rosberg et al. WACV 2023 Shiohara et al. CVPR 2023 Yoo et al. WACV 2023", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1927 }, { "text": "Yuan et al. ArXiv 2023 Zeng et al. AAAI 2023 Cui et al. CVPRW 2023 Jiang et al. CVPR 2023 Natsume et al. SIGGRAPH 2018 Face swapping: Huang et al. CVPR 2023 Shiohara et al. CVPR 2022 Zhao et al. CVPR 2021 Dong et al. ECCV 2022 Zhao et al. ICCV 2021 Ba et al. AAAI 2024 Nirkin et al. TPAMI 2022 General:", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1928 }, { "text": "Chen et al. CVPR 2022 Lin et al. CVPR 2024b Dong et al. CVPR 2023 Yan et al. CVPR 2024 Tan et al. CVPR 2024 Nguyen et al. CVPR 2024 Yao et al. ICCV 2023 Le et al. ICCV 2023 Larue et al. ICCV 2023 Yan et al. ICCV 2023 Sun et al. ICCV 2023 Ju et al. WACV 2024 Trinh et al. WACV 2021 Chen et al. NeurIPS 2022", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1929 }, { "text": "Tan et al. AAAI 2024 Yang et al. AAAI 2022 Kim et al. CVPRW 2021 Xu et al. WACVW 2023 Du et al. CIKM 2019 Lanzino et al. CVPRW 2024 Ciamarra et al. WACVW 2024 Hussain et al. WACV 2021 Jeong et al. WACV 2022 Jeong et al. AAAI 2022 Qian et al. ECCV 2020 Face synthesis: Hooda et al. WACV 2024 Tantaru et al. WACV 2024", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1930 }, { "text": "Fig. 1. A taxonomy of the state-of-the-art deepfake generation and detection methods. The methods are first divided according to the target task: generation versus detection. For each task, the methods are further divided into different kinds of architectures. For each architecture, we separate the methods based on the media types. Large groups are further divided according to the deepfake types presented in Section 3. References are", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1931 }, { "text": "clickable links to papers. Best viewed in color. majority of methods are based on convolutional neural networks (CNNs) [19], [21], [24], [25], transformers [32]\u2013 [34], or hybrid architectures that combine CNNs either with transformers [35]\u2013[37] or recurrent neural networks (RNNs) [38], [39]. For each type of architecture, we further divide the", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1932 }, { "text": "contributions with respect to the media types: image, video, audio or multimodal (audio-visual). Next, we present the main contributions in each category of articles included in the taxonomy. We further review existing datasets for deep- fake detection in image, video and audio. We then aggregate the reported performance levels of deepfake detectors on", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1933 }, { "text": "the most popular datasets, thus facilitating a direct com- parison of existing methods. In addition, we introduce a benchmark to test the generalization capacity of deepfake detectors to out-of-distribution content. Interestingly, we find that state-of-the-art deepfake detectors showcase poor generalization to realistic deepfakes generated by newer", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1934 }, { "text": "and more powerful generative models. Finally, we identify research gaps in current literature, proposing a series of future work directions that can lead to the development of better frameworks to detect deepfake media. In summary, our contribution is fourfold: \u2022 We conduct a comprehensive survey of deepfake", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1935 }, { "text": "generation and detection methods, comprising recent advancements in four domains: image, video, audio and multimodal. \u2022 We construct a taxonomy of deepfake generation and detection methods, categorizing research articles according to tasks, architectures and media types. \u2022 We collect and merge results reported on popu-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1936 }, { "text": "lar deepfake detection benchmarks, providing the means to easily assess the current performance levels of deepfake detectors. \u2022 We introduce a benchmark to test the out-of-domain generalization of deepfake detection models, show- ing that current detectors generally exhibit high per- formance drops on deepfakes generated by new and", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1937 }, { "text": "powerful generators. 2 RELATED SURVEYS Several attempts have been made to survey deepfake detec- tion and generation. In Table 1, we gather related surveys and illustrate the tasks, domains, methods and other aspects covered by the gathered surveys. Some surveys only cover the generation part [43], [45], while others are particularly", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1938 }, { "text": "focused on detection [40], [41], [44], [50]. Many surveys consider only one input media type, e.g. video [40], [42], [43], [45] or audio [44], [50]. There are a few surveys [41], [46], [47] that cover all media types (image, video, audio and multimodal), but only Masood et al. [46] and Patel et al. [47]", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1939 }, { "text": "address both detection and generation tasks. Although the surveys of Masood et al. [46] and Patel et al. [47] are compre- hensive, they do not cover the most recent developments, such as diffusion models and vision transformers. In summary, we find that existing surveys are either out- dated or limited in terms of coverage, including only specific", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1940 }, { "text": "tasks (generation or detection) or media types (image, audio or video). In contrast, we conduct an extensive survey of current literature, covering both generation and detection, as well as all deepfake media types. Moreover, we create a multi-level taxonomy to ease the navigation through the current deepfake literature, providing direct links to the", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1941 }, { "text": "IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. 50, NO. 1, NOVEMBER 2024 3 Target image Source identity Fake identity (a) Identity swapping. Target identity Source expression Fake expression (b) Facial expression swapping. Fake attribute (facial hair) Target identity (c) Facial attribute manipulation.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1942 }, { "text": "Fake video Target speech Source identity (d) Talking face synthesis. Target background Source identity Fake background (e) Background swapping. Fake speech Source voice I lost my wallet! Please send me some money. Target text (f) Text-to-speech synthesis. Morgan Freeman riding a unicorn on the red planet", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1943 }, { "text": "Fake image Text prompt (g) Text-to-image generation. Target speech Partially fake speech (h) Partial synthesis. Fig. 2. Deepfake types according to the general procedure used to synthesize the fake content. For deepfake types that apply to multiple domains, we provide the illustration for only one domain. Best viewed in color.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1944 }, { "text": "TABLE 1 Comparing our survey with related surveys in terms of the covered tasks, domains, methods and other aspects. There are at least three factors that differentiate our survey from each of the other surveys. Survey Task Domain Method Generation Detection Image Video Audio Multimodal GANs Diffusion", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1945 }, { "text": "CNNs RNNs Transformers Others Taxonomy Datasets New Benchmark Das et al. [40] \u2713 \u2713 \u2713\u2713 Heidari et al. [41] \u2713 \u2713\u2713\u2713\u2713 \u2713\u2713 \u2713 Kaur et al. [42] \u2713\u2713 \u2713 \u2713\u2713\u2713\u2713\u2713\u2713 \u2713\u2713 Lei et al. [43] \u2713 \u2713 \u2713\u2713 \u2713 \u2713 Li et al. [44] \u2713 \u2713 \u2713\u2713\u2713\u2713 \u2713\u2713 Li et al. [45] \u2713 \u2713 \u2713\u2713\u2713 \u2713\u2713 \u2713 Masood et al. [46] \u2713\u2713 \u2713\u2713\u2713\u2713\u2713 \u2713\u2713 \u2713 \u2713\u2713 Patel et al. [47] \u2713\u2713 \u2713\u2713\u2713\u2713\u2713 \u2713\u2713 \u2713 \u2713\u2713", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1946 }, { "text": "Pei et al. [48] \u2713\u2713 \u2713\u2713 \u2713\u2713\u2713\u2713\u2713\u2713 \u2713\u2713 Seow et al. [49] \u2713\u2713 \u2713\u2713 \u2713 \u2713\u2713 \u2713 \u2713 Yi et al. [50] \u2713 \u2713 \u2713\u2713\u2713\u2713 \u2713\u2713 Zhang [51] \u2713\u2713 \u2713\u2713\u2713 \u2713 \u2713\u2713 \u2713 \u2713 Ours \u2713\u2713 \u2713\u2713\u2713\u2713\u2713\u2713\u2713\u2713\u2713\u2713 \u2713\u2713\u2713 referenced papers. To our knowledge, our survey is the first to propose a novel benchmark to test the generalization capacity of deepfake detectors to out-of-distribution data.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1947 }, { "text": "3 DEEPFAKE TYPES To date, a number of alternative procedures have been employed to generate deepfakes. In order to simplify the task of producing realistic deepfakes, one commonly used procedure is to only alter a certain aspect of a media file, e.g. modifying the identity of a person in an existing video,", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1948 }, { "text": "or changing the emotion of a speech, while preserving the speech content and the speaker\u2019s identity. By employing recent and powerful generative models [3], [52], generating deepfake content from scratch has also become prevalent. We further categorize the deepfake content according to the procedure employed to obtain the respective deepfake", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1949 }, { "text": "type. We illustrate the identified categories in Figure 2 and present them in detail below. Interestingly, we identify deepfake categories that are domain-agnostic (which have been applied to all media types) and domain-specific (which have only been applied to a certain media type). Identity swapping. Deepfakes based on identity swapping", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1950 }, { "text": "imply replacing the identity information of a target person with that of a source person, while preserving identity- agnostic attributes, such as facial expressions. In the visual domain, this kind of deepfake is often referred to as face swapping, while in the audio domain, it is known as voice con-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1951 }, { "text": "version or voice swapping. Voice conversion seeks to change the timbre and prosody of a speaker with those of another speaker, while preserving the content of the speech. Expression/emotion swapping. In contrast to identity swapping, expression or emotion swapping involves al- tering the facial expression/emotion without changing the", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1952 }, { "text": "identity information. In the image domain, this task is known as facial expression swapping. In the video domain, the task is also known as face reenactment, and it implies altering the facial movement, which might often require facial motion capturing technology. In the audio domain, emotion swapping is the task of changing the emotion of", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1953 }, { "text": "IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. 50, NO. 1, NOVEMBER 2024 4 and video content via facial attribute manipulation implies changing certain semantic attributes of a target face, while maintaining the identity information. Some of the attributes that are usually altered are age, gender, skin color and hair.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1954 }, { "text": "Talking face synthesis. Talking face synthesis is perhaps the most complex procedure to obtain deepfake audio-video content, but also the most flexible procedure. The task seeks to generate an audio-video file of a talking face, where the source character is engaged in the act of speech. The gen- erated content is conditioned on some target information,", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1955 }, { "text": "provided in the form of text, audio, video or even mul- timodal content. The facial expressions, head movements, lip movements, speech emotions, and spoken content in the synthesized audio-video content are consistent with those of the source character. Background swapping. Deepfakes based on background", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1956 }, { "text": "swapping are generated by changing the background scene with a different one. In the visual domain, this involves segmenting the source person and blending this person in a new scene. In the audio domain, the background sound of the original recording is replaced with another background sound using audio editing technologies, while preserving", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1957 }, { "text": "speech content and the speaker\u2019s identity. Text-to-speech synthesis. Deepfake audio can be created with the help of a machine learning model for speech synthesis, starting from a piece of text. Current text-to- speech (TTS) synthesis models can produce natural speech and emulate the voice of a source identity.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1958 }, { "text": "Text-to-image/video generation. With the recent develop- ment of diffusion models, such as Stable Diffusion [3] and GLIDE [53], a new type of deepfake has emerged. Deepfake content can be easily generated by simply prompting a text- conditional diffusion model. All the necessary details, in- cluding the name of the source person, the facial attributes,", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1959 }, { "text": "the body pose and the actions can be specified through the prompt. This approach can be used to generate both images and videos. In text-to-speech synthesis, the input text is literally pronounced by the system, while in text- to-image/video generation, the prompt is rather interpreted by the system as a set of instructions.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1960 }, { "text": "Partial synthesis. As the name implies, partial synthesis involves changing only a part of an existing media file to create a deepfake. In the video domain, this kind of deepfake can be obtained by changing a subset of frames. In the audio domain, partial synthesis seeks to change only a subset of words in an utterance. The changes are performed", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1961 }, { "text": "in such a way that the target identity is maintained for the entire duration of the video or audio clip. 4 DEEPFAKE GENERATION In Figure 1, we first divide deepfake research by task, into deepfake generation methods and deepfake detection methods. We organize our literature review according to this", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1962 }, { "text": "split, discussing generative methods in the current section. We find that the most prominent approaches for deepfake generation are based on GANs or diffusion models. In some domains, such as video and audio, transformer-based methods are also very popular. Less frequently encountered methods are based on Variational Autoencoders (VAEs),", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1963 }, { "text": "Neural Radiance Fields (NeRF), 3D Morphable Models (3DMMs) and CNNs. Some models are only applied to specific media types, e.g. NeRF and 3DMMs are typically applied in the video domain. A number of studies use hy- brid models, combining GANs with VAEs on the one hand, or CNNS with RNNs on the other. We further structure", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1964 }, { "text": "our presentation according to the media type. For each media type, we divide the surveyed studies according to the underlying architectures. Moreover, we provide tuto- rials for the most important generative frameworks in the supplementary. 4.1 Image 4.1.1 GAN-based methods Face synthesis. Creating realistic faces is essential for deep-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1965 }, { "text": "fake generation, and GANs are widely employed to achieve this [5], [6], [54]\u2013[60]. Shen et al. [54] introduce a third model into the traditional adversarial framework, tasked with de- termining whether the generated images retain the identity from a reference image. This approach enables the method", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1966 }, { "text": "to perform conditional generation. Conditional generation is also the focus of Xu et al. [60]. Their approach synthesizes high-quality 3D heads with control over the camera poses and other facial attributes. StyleGAN [56]\u2013[58] improves the quality of the synthesized images by changing the generator architecture, and leveraging a mapping network", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1967 }, { "text": "to map the usual Gaussian vector to an intermediary latent space. The generative network adopts these new latents at different scales in the architecture via AdaIN [61] layers. Fu et al. [59] demonstrate that StyleGAN is also effective for generating full body images. Sauer et al. [5] extend the StyleGAN model, presenting a method that leverages Pro-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1968 }, { "text": "jected GAN training [62], progressive growing and classifier guidance [63], unlocking image synthesis at a resolution of 1024\u00d71024. Different from StyleGAN, Esser et al. [6] utilize GANs to learn a perceptually rich codebook, representing images as a sequence of codebook entries. This approach enables the use of transformers for training.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1969 }, { "text": "Face swapping. One of the most widely-used methods for generating deepfakes is face swapping. This technique involves replacing the face in a target image with that of another individual, sourced from a different image. The key challenge lies in seamlessly integrating the source face, while maintaining non-identity-specific attributes, such", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1970 }, { "text": "as facial expressions and lighting conditions. Thanks to their well-known capacity of generating realistic images, GANs [7]\u2013[10], [27], [55], [64]\u2013[69] are widely adopted in face swapping frameworks. In GAN-based face swapping pipelines, the generator is usually conditioned on identity information from the", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1971 }, { "text": "source image and the attributes extracted from the target image [9], [10], [65], [66], [68]\u2013[73]. The work of Bao et al. [65] is one of the earliest contributions in this direction. The authors employ a face recognition model to extract an identity embedding from the source image. The attributes of the target image are extracted by a neural network", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1972 }, { "text": "trained to minimize the Euclidean distance between the target and generated images, applying a lower weight when the identities in the target and source images differ. The concatenated representations are processed by a generator that is trained in an adversarial setting. Li et al. [69] ad- vance the previous framework by introducing a multi-level", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1973 }, { "text": "IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. 50, NO. 1, NOVEMBER 2024 5 providing a more detailed representation of the target image than the approach of Bao et al. [65]. Similarly, Chen et al. [66] propose the ID Injection Module, which integrates iden- tity information through Adaptive Instance Normalization", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1974 }, { "text": "(AdaIN) [61] layers into the target image features. More recent works [9], [10], [68], [70], [71], [74] increase the quality and quantity of the conditional identity information. Kim et al. [9] enforce smoothness to the identity encoding space through contrastive learning. Cao et al. [74] and Cui et al. [72]", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1975 }, { "text": "explore the effectiveness of the transformer architecture for identity embedding in face swapping. Shiohara et al. [10] improve the embeddings by introducing BlendFace, a face encoder model trained with synthetic face images featuring swapped attributes. Rosberg et al. [68] leverage the feature maps provided by multiple layers of the face encoder to", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1976 }, { "text": "better represent the identity. Zeng et al. [70] show that a masked autoencoder (MAE) [75], pre-trained on a large- scale face dataset, is an effective encoder for face swapping. Yoo et al. [71] present the Triplet Adaptive Normalization block to integrate the pose and identity features within their", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1977 }, { "text": "generator. Slightly distinct from earlier studies, another line of re- search [7], [8], [27], [67], [76]\u2013[78] explores the manipulation of identity and attribute features within the latent space of the generator. Natsume et al. [27] create the latent space of the generator by merging the outputs of two encoders. One", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1978 }, { "text": "of the encoders is responsible for identity information, and the other for attributes. Two encoders are also utilized by Ren et al. [67] to separately learn embeddings for facial non- identity and non-facial attributes. This separation eliminates the need for skip connections, preventing identity leakage", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1979 }, { "text": "from the target image. Gao et al. [76] perform the disentan- glement between identity and attributes through their novel Informative Identity Bottleneck layers that are included in a frozen face recognition model. Zhu et al. [77] train a model to perform GAN inversion and obtain the latent code for a", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1980 }, { "text": "given image. Subsequently, they employ another model to integrate the attributes from the target image into the latent code of the source face. The resulting latent code is fed into StyleGAN2 [57] to generate the swapped image. Similarly, Li et al. [7] employ learnable GAN inversion, but in their case, they leverage the latent space of a 3D GAN [79] to", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1981 }, { "text": "synthesize multi-view swapped images. The latent space of StyleGAN is further exploited for face swapping by Liu et al. [8], where the GAN inversion is extended at region level through the use of facial semantic masks. Jiang et al. [78] introduce identity preserving semantic bases (StyleIPSB) for StyleGAN [56]. Their approach identifies direction vectors", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1982 }, { "text": "in the latent space that modify attributes of generated im- ages, while maintaining the identity of the subject. GANs are also used in face swapping for the purpose of fixing the swapped image and making it more realistic [80]\u2013 [82]. Specifically, Moniz et al. [80] and Sun et al. [81] employ GANs to perform the blending of the source face in the", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1983 }, { "text": "target image, leveraging inpainting pipelines or frameworks such as CycleGAN [83]. Chen et al. [82] take a step further and present a method to correct deepfake images perturbed with adversarial attacks. Face editing. Altering facial attributes such as age, gender, hair color or pose can be used to create counterfeit content.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1984 }, { "text": "GANs support this kind of applications, either via image- to-image translation between different domains [83]\u2013[86] or via latent code manipulation [15], [87], [88]. CycleGAN [83] was first proposed for image-to-image translation between two domains. The method comprises two generators, one for each of the two domains. The primary contribution of", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1985 }, { "text": "Zhu et al. [83] is the introduction of the cycle-consistency loss, which ensures that the pipeline can reconstruct the original image, after translating it from one domain to the other and back. The main limitation of CycleGAN is its ability to handle only two domains. StarGAN [84], [85] addresses this limitation and supports image translation", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1986 }, { "text": "from multiple domains. The model allows this feature by including an additional condition as input, along with the conditional image. Hsu et al. [86] employ a dual-generator approach for image-to-image translation. The first generator produces a landmark image matching the pose of a reference image, and the second uses this landmark to recreate a", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1987 }, { "text": "source identity in the specified pose. Different from these approaches, other works [15], [87], [88] harness the latent space of StyleGAN. Tov et al. [87] study the latent space of StyleGAN and design an encoder for inversion, which is suitable for image editing. Similarly, Suwa\u0142a et al. [88] design a plugin for the latent space of StyleGAN. This", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1988 }, { "text": "plugin disentangles the latent codes into attribute and non- attribute features, allowing attribute manipulation for facial editing. Bounareli et al. [15] project an identity image in the latent space of StyleGAN, and then, they harness pose and appearance encoders to create offsets, allowing the generator to change the pose of the identity latent.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1989 }, { "text": "4.1.2 Diffusion-based methods Text-to-image. Diffusion models are effectively applied in text-to-image generation [2]\u2013[4], [89], utilizing large lan- guage models to encode textual descriptions that guide image creation. This capability allows users to generate counterfeit content featuring public figures simply by in-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1990 }, { "text": "cluding their names in the text description used as input for generation. One of the most popular methods for text-to- image generation is Stable Diffusion [3], which leverages the latent space of a vector quantized (VQ) GAN [6] to perform the diffusion processes. SDXL [4] scales up the Stable Diffusion architecture, improving the quality and text", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1991 }, { "text": "fidelity of the generated images. Personalized generation. Although text-to-image diffusion models allow deepfake content generation of public figures, some results do not accurately replicate the identity of the person. Thus, these models might have limited appli- cation in deepfake generation. However, there is another", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1992 }, { "text": "direction of research [11]\u2013[13], [29], [30], [90]\u2013[110] focused on generating images that contain a specific identity or concept depicted in an image or a set of images given as input. Such methods are more likely to be employed in deepfake generation. We can distinguish these contributions into two main approaches, those that perform test-time fine-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1993 }, { "text": "tuning [90], [94], [97], [100], [102], [111]\u2013[113] and those that leverage large-scale datasets and learn how to incorporate the additional images offline [11]\u2013[13], [29], [30], [91]\u2013[93], [95], [96], [98], [99], [102]\u2013[110]. Test-time fine-tuning approaches use different compo- nents to integrate and learn the new identity. Some works", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1994 }, { "text": "IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. 50, NO. 1, NOVEMBER 2024 6 embed it [97], [100], [111]. Other approaches [94] either use low-rank adaptation (LoRA) [114] or directly fine-tune the weights of the denoising network [90], [102], [112], [113]. Overall, test-time fine-tuning methods yield impressive re-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1995 }, { "text": "sults in terms of identity preservation, but their main disad- vantage is the expensive optimization, which significantly increases the generation time. To this end, many works address the efficiency issue, to some extent. For example, Chen et al. [113] try to incorporate the knowledge of mul- tiple subject-specific models into a single model. Ruiz et", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1996 }, { "text": "al. [112] leverage a HyperNetwork architecture to predict the network weights from a face image. Subsequently, they use these weights as a starting point for test-time fine-tuning, reaching faster convergence than previous work [90]. De- spite these advancements, test-time fine-tuning methods still suffer from high generation times.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1997 }, { "text": "In contrast to test-time fine-tuning methods, the ap- proaches that harness offline training [11]\u2013[13], [29], [30], [91]\u2013[93], [95], [96], [98], [99], [102]\u2013[109] are faster in terms of generation time, but their primary issue is identity preser- vation. Therefore, solving the latter problem constitutes", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1998 }, { "text": "the priority of these works. Zhao et al. [12] propose an identity preservation loss for which they construct a better estimation of the original image given the predicted noise, at training time. The same idea is studied by Liu et al. [13], who improve the estimation even further. Peng et al. [92] employ an identity loss, but only for certain noise levels.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 1999 }, { "text": "Other methods [11], inspired by the GAN literature, em- ploy the ArcFace model as identity embedding extractor for better identity representations. Similarly, Li et al. [109] improve representations by stacking multiple embeddings when multiple images are available. Lastly, Wang et al. [30] decouple the generation of background and identity-related", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2000 }, { "text": "content by training two separate denoising networks, out of which only one knows to generate images of a given person. Tools. The most popular tools for personalized genera- tion are LoRA-based variants of Stable Diffusion [3] and SDXL [4], that are specialized on particular public personal- ities, e.g. Elon Musk3 or Alan Turing4. Different from these", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2001 }, { "text": "options, another powerful tool is Midjourney5. In contrast to the LoRA-based methods, Midjourney is not popular for personalized generation, but for text-to-image synthesis. However, given the quality of its generative results, Mid- journey is a popular tool for creating counterfeit images. 4.1.3 Other methods", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2002 }, { "text": "Unlike previous methods centered around generative mod- els, some approaches rely on alternative techniques. Bitouk et al. [115] identify the closest match in terms of lighting and pose from a large set of face images, and perform face replacement using key point alignment. Korshunova et al. [116] use a multi-resolution CNN in the VGG feature", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2003 }, { "text": "space, aligning target and source images to minimize cosine distance between corresponding patches. Wang et al. [117] propose an encoder-decoder architecture with a 3D identity extractor and a Semantic Facial Fusion module to enhance resolution and preserve identity. In contrast, other works 3. https://civitai.com/models/603798/elon-musk-sdxl", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2004 }, { "text": "4. https://civitai.com/models/796450/alan-turing-mathematical- flux 5. https://www.midjourney.com/ combine generative methods to produce higher-quality im- ages. Bao et al. [118] introduce CVAE-GAN, a method which combines VAEs with GANs. The generator and the encoder are trained with an adversarial objective, but also with a", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2005 }, { "text": "mean feature matching objective and a pixel-wise recon- struction loss, respectively. Li et al. [119] merge diffusion models and GANs by representing identity in Stable Diffu- sion through the latent space of StyleGAN, integrating latent codes into the U-Net via cross-attention layers. 4.2 Video 4.2.1", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2006 }, { "text": "GAN-based methods The early works for generating deepfake videos employ conventional GAN models for face swapping and reen- actment, which are applied frame by frame [14], [120]. Nevertheless, these methods are usually part of more com- plex frameworks which have zero-shot capabilities, either involving more steps [14] or enhanced architectures [120].", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2007 }, { "text": "Gao et al. [121] introduce a face reenactment GAN, focusing on generating videos of talking heads. The facial landmarks, expressions and head poses are extracted from both source and target frames to fit a face 3DMM and obtain predefined keypoints. To depart from the conventional paradigm and improve", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2008 }, { "text": "the video generation using GANs, subsequent works lever- age the latent space of StyleGAN2 [57]. A consistent number of methods divide the latent space in which they operate into two: one for content and one for motion [28], [122], [123]. While some utilize an RNN for sampling the motion trajectory [122], [123] and employ two discriminators, one", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2009 }, { "text": "for individual frames and one for the video sequence, Sko- rokhodov et al. [28] compute the motion embeddings with 1D convolutional layers and use only one video discrimina- tor. Oorloff et al. [16] take a different approach by encoding both source and target frames, fusing their latent representa-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2010 }, { "text": "tions, then generating the output frame, while also utilizing multiple latent spaces [124], [125]. Yu et al. [126] treat videos as continuous-time signals and, with the help of an Implicit Neural Representation [127], they map an input signal (pixel coordinates and time) to RGB values in order to generate the", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2011 }, { "text": "corresponding video. Brook et al. [128] propose to generate multiple consecutive frames in low-resolution, and then increase their resolution with a super-resolution network. This approach ensures that training long video sequences is feasible. 4.2.2 Diffusion-based methods Latent diffusion models [3] use a cross-attention mecha-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2012 }, { "text": "nism that facilitates conditioning diffusion models for im- age generation. However, the main challenge in generating deepfake videos with diffusion models is employing a con- ditioning mechanism, while achieving temporal cohesion. The studies of Ho et al. [129] and Blattman et al. [130] rep- resent the stepping stones in adopting diffusion models for", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2013 }, { "text": "video generation. Their methods extend diffusion models in several ways. The architectural changes applied on the U-Net mainly consist of replacing 2D convolutions with 3D convolutions, and appending additional self-attention lay- ers for temporal attention. In a subsequent work, Blattmann et al. [131] demonstrate the benefits of using a large curated", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2014 }, { "text": "IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. 50, NO. 1, NOVEMBER 2024 7 dataset for training a video generator. Wu et al. [132] in- troduce a one-shot method for editing a video given a text prompt. Inspired by Ho et al. [129], a text-to-image diffusion model is extended to an additional dimension (time) by Wu", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2015 }, { "text": "et al. [132], where the added self-attention layers operate on the current frame and the previous two frames. Newer diffusion-based video generation methods [133]\u2013 [135] depart from the U-Net architecture and adopt a transformer-based one, namely ViT, which provides an in- nate mechanism for both spatial and temporal attention.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2016 }, { "text": "This allows longer videos to be generated. Additionally, Guo et al. [134] introduce a plug-and-play module that can be integrated into a text-to-image diffusion model to induce the ability to generate videos. Inspired by this module, Wang et al. [136] present a method for text-to- video generation composed of several stages, in which a", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2017 }, { "text": "ControlNet is applied to improve guidance. Different from previous studies employing Stable Diffusion as the base model, Singer et al. [137] ground their work on DALLE-2 [138], while Ho et al. [139] utilize Imagen [89]. Nevertheless, similar architectural changes are implemented, where the network is extended to support the temporal dimension.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2018 }, { "text": "An important line of research is represented by portrait animation, in which a video is generated from a source frame and various conditional inputs. Most works in this area [140]\u2013[143] aim to apply a sequence of facial expres- sions over the image. Two different approaches are used to condition the diffusion model. One is based on intermediate", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2019 }, { "text": "representations of the facial expressions, such as facial key- points [140], [141], [143], and the other is based on directly encoding the frames containing the target facial movements [142]. Currently, the ability of the video generation methods based on diffusion modeling is not satisfactory, often requir-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2020 }, { "text": "ing quality enhancements at a later stage in the pipeline. For example, super-resolution models are sometimes employed to increase video resolution [136], [137], [139], while the frame rate is usually increased through frame interpolation [131], [136], [137], [141]. 4.2.3 Other methods Transformers represent the most popular architectural", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2021 }, { "text": "choice for video generation. For instance, Rochow et al. [144] leverage cross-attention blocks to guide the generation (us- ing encoded facial keypoints and expressions), while Ville- gas et al. [145] apply the attention mechanism on frames to generate longer and coherent videos. An alternative choice for video generation consists of", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2022 }, { "text": "employing some VQ autoencoder, either variational [146] or standard [147]. Similar to diffusion models, the generation process is carried out in the latent space of the autoencoder, which is lower dimensional. Within this vector space, a transformer is used to generate video tokens [147]\u2013[150]. Unlike other related approaches, Jiang et al. [149] carefully", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2023 }, { "text": "design the latent space such that it is decomposed into an appearance and a pose representation, respectively. A few methods harness the 3D space for face reenact- ment. In this context, warping is often employed, which involves computing a flow field between the source frame and the driving frame, then applying it on the former", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2024 }, { "text": "frame. In the same context, NeRF models [151] are used to generate novel views of 3D face models. For example, Thies et al. [152] obtain a coarse 3D representation from the source frame using a traditional graphics pipeline, and then feed it to a neural network to obtain a neural texture, a high- dimensional embedding space, from which a Deferred Neu-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2025 }, { "text": "ral Renderer (based on U-Net) generates the target image. Thies et al. [153] and Yang et al. [154] synthesize faces by applying a deformation transfer between two 3DMM-based intermediate representations of the source and driving video frames, the mouth being further refined through warping. Zhang et al. [155] also apply warping based on dense land-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2026 }, { "text": "marks, while Li et al. [156] combine warping with NeRF. Finally, to increase the performance, some works adopt pre- training strategies that involve masking the input and then reconstruct the signal [145], [147]. 4.3 Audio 4.3.1 GAN/VAE-based methods A number of text-to-speech models employ popular gen-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2027 }, { "text": "erative frameworks, such GANs and VAEs, either alone or in combination with more recent developments in the field. Kim et al. [157] propose an end-to-end TTS framework that augments variational inference with normalizing flows and uses an adversarial training procedure to enhance the representation potential. The method of Casanova et al. [158]", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2028 }, { "text": "constructs on the previous model, introducing new proce- dures, such as the concatenation of language embeddings with the ones of the input characters to allow training in a multilingual fashion. In [159], the authors introduce new modules to develop NaturalSpeech, another VAE-based TTS. A differentiable", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2029 }, { "text": "durator is used to improve the duration prediction, a mem- ory mechanism simplifies the waveform reconstruction, and a bidirectional prior/posterior module improves the prior from text, while simplifying the posterior from speech. Lee et al. [160] present a GAN-based vocoder that im- proves the generator by introducing anti-aliased feature", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2030 }, { "text": "representation and periodic non-linearities, delivering state- of-the-art results and robustness for out-of-distribution sce- narios, such as novel languages and speakers. 4.3.2 Transformer-based methods A few recent methods employ transformers to obtain com- petitive generation performance. FastSpeech [161] intro-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2031 }, { "text": "duces a transformer-based model that speeds up speech synthesis by parallelizing Mel-spectrogram generation through a feed-forward architecture. A length regulator is used to match the length of the hidden states with the length of the Mel-spectrograms, and a duration predictor provides the duration for the phonemes. Jiang et al. [162]", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2032 }, { "text": "reuse the length regulator and the duration predictor from FastSpeech, adding separate modules for content, timbre and prosody modeling. A global timbre encoder is used to extract a global timbre vector, while a latent code language model fits the prosody distribution. Wang et al. [163] proposed VALL-E, a framework that", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2033 }, { "text": "IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. 50, NO. 1, NOVEMBER 2024 8 information and content. SPEAR-TTS [164] removes the necessity to supply the transcripts of audio prompts by de- coupling the generation of semantic tokens and the acoustic tokens. Yang et al. [165] introduce UniAudio, a hierarchical trans-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2034 }, { "text": "former framework that learns both inter-frame and intra- frame correlations separately, reducing the computational complexity. It supports the generation of multiple types of audio by employing LLM-style next token prediction and tokenization via a universal neural codec. 4.3.3 Diffusion-based methods", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2035 }, { "text": "Following the success of diffusion models in vision [1], several generation methods adopted the diffusion mod- eling framework to generate deepfake audio. Huang et al. [166] present FastDiff-TTS, a conditional diffusion model that follows the architectural design proposed by Ren et al. [167]. The authors employ time-aware location variable", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2036 }, { "text": "convolutions for long-term dependency modeling and a noise schedule predictor for sampling acceleration. ProDiff [168] is another framework with an architecture inspired by Ren et al. [167], which uses a denoising model with a parametrization that directly predicts the clean data, halving the number of diffusion steps via knowledge distillation.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2037 }, { "text": "The audio encoder/decoder, the phoneme encoder and the duration and pitch predictors proposed by Tan et al. [159] are reused in NaturalSpeech 2 [169], alongside a dif- fusion model that learns to predict latent representations conditioned on the input text. To promote zero-shot gener- ation, a speech prompting mechanism helps the diffusion", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2038 }, { "text": "model and the duration and pitch predictors to follow prosody, style and speaker identity from the supplied audio prompt. The encoder/decoder and the duration predictor are further used in NaturalSpeech 3 [170], where, in contrast to previous studies [159], [169], each of the following speech attributes are independently generated by a novel factorized", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2039 }, { "text": "diffusion model: duration, content, prosody and acoustic details. Du et al. [171] introduce UniCATS, a framework capable of performing speech editing tasks, where speech is synthe- sized by taking into account both preceding and following contexts. UniCATS can achieve this with a contextual VQ- diffusion-based acoustic model and a contextual vocoder.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2040 }, { "text": "Yang et al. [172] aim to solve learning problems specific to expressive TTS, a novel task that tries to control the speaking style of the synthesized speech. The proposed framework uses RoBERTa [173] to extract the style representation from a natural language prompt. 4.4 Multimodal 4.4.1 Transformer-based methods", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2041 }, { "text": "Recent advancements in talking face generation focus on improving the synchronization of facial movements with speech, while maintaining natural motion and emotional consistency [174]\u2013[177]. These approaches address chal- lenges such as lip-sync accuracy [175]\u2013[177], motion stability [174], [177] and speaker-specific styles [174], [175], aiming", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2042 }, { "text": "for realistic human-video synthesis. Jang et al. [174] intro- duce a system that combines talking face generation with TTS, addressing the challenge of generating natural head poses and maintaining consistent speech patterns even with varying facial motions. Their approach leverages a motion sampler and a conditioning method to ensure fluidity in", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2043 }, { "text": "both aspects. Building on the idea of synchronizing audio and visual elements, Cheng et al. [175] propose VideoReTalk- ing, a method designed to edit real-world talking head videos for perfect lip-sync and emotional consistency. In a similar fashion, Wang et al. [176] develop a one-shot talking face generation framework. They introduce an audio-visual", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2044 }, { "text": "correlation transformer, which improves lip-sync accuracy by mapping audio to dense motion fields through phoneme and keypoint representations. Addressing another challenge in the field, Ling et al. [177] tackle the issue of lip motion jit- ter in speech-driven talking face generation. Their solution,", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2045 }, { "text": "StableFace, identifies key problems such as noise in the 3D face representation and mismatches between training and inference stages. To address emotion-agnostic talking head generation, Gan et al. [178] propose emotional adaptation for audio-driven talking-head. The method enhances emotion- agnostic talking-head models by adding three lightweight", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2046 }, { "text": "adaptations: deep emotional prompts, an emotional defor- mation network, and an emotional adaptation module. 4.4.2 Diffusion-based methods Several frameworks are designed to generate high-quality audio-driven portrait animations, aiming to achieve real- ism and synchronization [18], [31], [179]\u2013[183]. AniPortrait", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2047 }, { "text": "EchoMimic [181] offers another solution by integrating au- dio and facial landmarks using a denoising U-Net archi- tecture, which stabilizes and enhances the natural flow of portrait videos. Similarly, Stypu\u0142kowski et al. [31] propose an autoregressive diffusion model to achieve realistic talking heads with smooth, expressive movements and accurate lip-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2048 }, { "text": "sync. Xu et al. [18] also use a diffusion-based framework to improve lip-sync accuracy and motion diversity, employing a hierarchical audio-driven visual synthesis module. In a similar direction, VASA [182] produces talking faces, cap- turing synchronized lip movements and dynamic expres- sions using a diffusion-based model in a latent facial space,", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2049 }, { "text": "enabling real-time interactions with high realism. Distinctly, EMO [183] generates expressive talking head videos with- out relying on 3D models, excelling in natural transitions and seamless identity preservation. 4.4.3 Other methods Recent advancements in talking head generation leverage different models, such as NeRF [184], [185], GANs [186]\u2013", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2050 }, { "text": "IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. 50, NO. 1, NOVEMBER 2024 9 videos with arbitrary speech audio. By improving audio-lip synchronization using pitch contour analysis and incorpo- rating a fast motion-to-video renderer, GeneFace++ offers a robust and efficient solution.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2051 }, { "text": "In the context of GANs, Text2Video [186] presents an approach for synthesizing videos directly from text, reduc- ing reliance on audio-driven models. Using a phoneme- pose dictionary and a GAN-based architecture, the method achieves high-quality video synthesis with just one minute of training data. HeadGAN [187] is developed for head", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2052 }, { "text": "reenactment and editing from a single reference image. It integrates 3DMMs for real-time reenactment at approxi- mately 20 FPS. Furthermore, it incorporates audio features for enhanced mouth movement accuracy. Wang et al. [188] introduce TalkLip, a speech-to-lip generation model that enhances lip-speech intelligibility by incorporating a lip-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2053 }, { "text": "reading expert to penalize incorrect outputs. For gesture generation, RNN-based frameworks such as hierarchical audio-to-gesture (HA2G) [189] introduce ways to generate co-speech gestures. HA2G extracts multi-level audio features using a hierarchical audio learner. Another RNN-based model [190] offers a real-time pipeline to gen-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2054 }, { "text": "erate personalized photorealistic talking-head animations. This model operates at over 30 FPS and follows a three- stage process: extracting deep audio features, predicting facial dynamics and head motions with an auto-regressive model, and rendering high-fidelity faces through image-to- image translation. Some RNN-based methods [194], [195]", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2055 }, { "text": "rely on audio-visual cues for realistic face synthesis. Peng et al. [194] present a speech-driven 3D face animation model that separates speech content and emotion using an Emotion Disentangling Encoder, while Zhong et al. [195] introduce a two-stage framework for audio-driven person-generic talk- ing face video generation. Both approaches apply RNNs", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2056 }, { "text": "on top of CNN features. The CNN-based DisCoHead [193] offers an unsupervised approach to disentangling head mo- tion from facial expressions. By applying geometric transfor- mations to isolate head motion and using speech audio for facial expressions, DisCoHead efficiently generates realistic talking heads.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2057 }, { "text": "5 DEEPFAKE DETECTION The second part of our taxonomy illustrated in Figure 1 comprises deepfake detection methods. The taxonomy clearly indicates that most deepfake detectors are based on CNN architectures. However, with the recent advent of vision and audio transformers, a large body of work on deepfake detection is now based on multi-head attention.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2058 }, { "text": "To boost detection performance, a considerable number of studies employ hybrid models, combining CNNs with transformers or RNNs, respectively. Less prevalent architec- tures in deepfake detection are graph and recurrent neu- ral networks. We organize our subsequent presentation of deepfake detection methods according to the input domain.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2059 }, { "text": "The reviewed articles are further separated according to the employed architectures. 5.1 Image 5.1.1 CNN-based methods Convolutional nets are the most prevalent type of architec- ture for deepfake detection [19]\u2013[22], [196]\u2013[223]. The detec- tion task is commonly formalized as a binary classification,", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2060 }, { "text": "where a CNN backbone is used as feature extractor. The most frequently chosen backbones in this line of work are EfficientNet [224], XceptionNet [225] and ResNet [226]. The main direction of research for deepfake detection focuses on developing methods that generalize well across different types of manipulations [20], [21], [196]\u2013[206], [208],", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2061 }, { "text": "and complexity of forgeries, but different from Chen et al. [20], they achieve this through latent space manipula- tions. Other studies [21], [198], [200], [202], [203], [206], [208] try to identify the common artifacts or features for different types of forgeries. Thus, some methods [21], [198], [206], [208] base their solution on local artifact detection", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2062 }, { "text": "and less on the overall identity. Yan et al. [200] propose an explicit disentanglement approach using multi-task learning to analyze image information, allowing the detection of features that are shared across various types of forgeries. Some studies [202], [203], [223], [228], [229] indicate that the", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2063 }, { "text": "generalization capabilities of deepfake detectors can also be improved by leveraging artifacts in the frequency domain. All these advancements in the generalization of deepfake detectors are validated by Yao et al. [199], who demonstrate that detectors modeling low-order interactions exhibit su- perior generalization capabilities. In contrast to previous", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2064 }, { "text": "works, some deepfake detection methods [19], [207], [214], [216] are specialized in identifying specific forgery tech- niques. Huang et al. [207] and Shiohara et al. [19] focus on face swapping. Huang et al. [207] argue that a face-swapped image contains information about the identity present in the target image. Their method builds a face recognition", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2065 }, { "text": "model to detect this identity from a face-swapped image, and leverages the difference between its embedding and the embedding of the source identity to detect deepfakes. Shiohara et al. [19] improve face-swapping detection by generating more challenging examples for a face-swapping detector. In this regard, they use face-swapping pipelines", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2066 }, { "text": "where the target and source images are of the same identity and closely resemble each other in terms of face position and other non-identity attributes. In contrast, Hooda et al. [214] and Tantaru et al. [216] tackle the detection of forgeries in images generated with diffusion models. Hooda et al. [214]", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2067 }, { "text": "detect deepfakes using an ensemble in which the models are using disjoint parts of the input features, aligning with the aforementioned finding of Yao et al. [199]. Other studies explore deepfake detection methods by examining their fairness [22], [215], level of explainabil- ity [209], [217], or level of vulnerability to adversarial at-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2068 }, { "text": "tacks [230]. Ju et al. [215] propose the first approach to tackle fairness in deepfake detectors. Their method uses groups of people specified by the user and ensures that the loss of these groups is similar to each other. The method of Ju et al. [215] performs well when tested on the same type of forgery as in the training set. Lin et al. [22] extend the", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2069 }, { "text": "IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. 50, NO. 1, NOVEMBER 2024 10 graphic and forgery specific features. These features are then combined and used in a fairness loss function, which aims to ensure equal importance across different demographic groups. In terms of explainability, Dong et al. [209] try to", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2070 }, { "text": "identify the visual concepts that are relevant for deepfake detectors. Their findings indicate that the features specific to source and target images are, in general, ignored. The focus of the detection models is on visual artifacts. Trinh et al. [217] reinforce this finding by showing that, in addition", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2071 }, { "text": "to visual artifacts, temporal artifacts also serve as evidence for detectors. Hussain et al. [230] show that, despite the progress of deepfake detectors, these models are susceptible to adversarial attacks and future works need to address this drawback. 5.1.2 GCN-based methods As stated before, Yao et al. [199] demonstrate that deepfake", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2072 }, { "text": "detectors with strong generalization capabilities tend to model low-order interactions. Due to their ability to capture such relationships, some studies employ graph convolu- tional networks (GCNs) for deepfake detection [231]\u2013[233]. Yang et al. [231] design their graphs with vertices represent- ing features of facial regions and edges capturing the corre-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2073 }, { "text": "lations between these regions. They also introduce a mask- ing strategy that removes edges based on their weights. These graphs are then given as input to a GCN, which extracts features for a binary classifier. Wang et al. [232] propose a similar method, but along with the spatial domain features, they also include frequency domain information.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2074 }, { "text": "5.1.3 Transformer-based methods State-of-the-art results in a broad range of computer vision tasks are achieved by transformer architectures [234]. As a result, these architectures are often adopted for deepfake detection [32], [33], [235]. Aghasanli et al. [235] propose a direct application of transformers for deepfake detection,", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2075 }, { "text": "where the model is used as feature extractor for a binary classifier. The method of Dong et al. [33] is similar to that of Huang et al. [207], as both approaches harness discrepancies between the explicit identity depicted in the image and the one given by the outer region of the face. However, in contrast to Huang et al. [207], Dong et al. [33] use the", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2076 }, { "text": "same model for both identities and differentiate between the two regions with two additional tokens. Wang et al. [32] use a multi-scale transformer which operates on patches of different sizes to extract features for deepfake detection. 5.1.4 Hybrid architectures A natural strategy for improving deepfake detection is to", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2077 }, { "text": "combine the aforementioned architectures in a joint pipeline. The primary research focus is on combining transformer and CNN architectures [37], [236], [237], though the combination of GANs and CNNs [238] is also explored. For instance, Jeong et al. [238] train a GAN to generate perturbation maps, which are added to both real and fake images to minimize", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2078 }, { "text": "differences at the frequency level. They argue that using this augmented data to train a CNN-based classifier prevents overfitting to method-specific frequency artifacts, thereby improving the generalization. Kamat et al. [237] focus on exploring different techniques of combining CNN-based and transformer-based feature ex-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2079 }, { "text": "tractors for deepfake detection. Shao et al. [37] and Hong et al. [236] harness the CNN and transformer combination for a slightly different problem. Their goal is to determine the sequence of facial manipulations used to create a fake im- age, because deepfakes are frequently created with several manipulation steps. Both methods [37], [236] rely on a CNN", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2080 }, { "text": "backbone for feature extraction followed by a transformer that returns the sequence of manipulated regions. 5.1.5 Traditional machine learning methods The earliest approaches to deepfake detection examine the effectiveness of classical machine learning algorithms [239], [240]. Guarnera et al. [240] conjecture that transposed con-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2081 }, { "text": "volutional layers within GANs generate local pixel corre- lations. Leveraging this, they design an algorithm to ex- tract local features from images, which are then passed to machine learning algorithms such as SVM and k-NN for detection. Lugstein et al. [239] also employ an SVM, but focus on extracting features from the photo response non-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2082 }, { "text": "uniformity (PRNU) signal. 5.2 Video 5.2.1 CNN-based methods Most deepfake video detection methods are based on plain convolutional models. In the majority of works, 3D convo- lutions are applied to extract spatio-temporal features from a whole video sequence. Nevertheless, 2D convolutions are also used to extract salient spatial features from individual", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2083 }, { "text": "frames, and then, the results are aggregated to make the final prediction. Agarwal et al. [241] extract facial features (both static and temporal), and then compare them with a reference set, for each biometric source data, to identify a similar data point, whose label is used for prediction. Simi-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2084 }, { "text": "larly, Cozzolino et al. [242] compare the extracted biometric features from the input video to those of a pristine video. Some works propose to capture temporal inconsistencies in the video. This is either achieved from successive frames, with some specialized sequential convolutional blocks [243],", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2085 }, { "text": "or through a hierarchical framework from both local (frame) and global (video snippet) perspectives that can differentiate between real and fake videos [244], [245]. Based on this objective, some papers focus only on inconsistencies of specific aspects of the face. Haliassos et al. [246], [247] study", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2086 }, { "text": "mouth movements and propose to learn spatio-temporal representations of mouth motion via two-stage frameworks. Demir et al. [248] magnify the motion of the face and then classify the videos, while also identifying the source generation method. 5.2.2 Hybrid CNN and RNN architectures To obtain a prediction based on the temporal dimension,", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2087 }, { "text": "many detection methods employ a recurrent network to aggregate the latent features extracted by CNNs. Multiple variants of recurrent architectures are used, such as simple RNNs [38], [249], gated recurrent units (GRUs) [39], [250], [251] and Long Short-Term Memory (LSTM) networks [252], [253]. Unlike the rest, Masi et al. [253] use two branches, each", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2088 }, { "text": "IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. 50, NO. 1, NOVEMBER 2024 11 and employ ArcFace [254] for a better representation of the backbone features. 5.2.3 Other hybrid architectures Aside from combining CNNs and RNNs, some attempts try to fuse other types of neural networks. An important", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2089 }, { "text": "category is represented by the integration of attention [255]\u2013 [257]. Bonettini et al. [255] integrate an attention mecha- nism into each network in an ensemble of CNNs. Wang et al. [256] extract noise features from the face crop, as well as a background crop, and feed them into a multi- head attention module. Furthermore, these works adopt", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2090 }, { "text": "the contrastive learning paradigm in their deepfake video detectors. Different from previous methods, Coccomini et al. [258] combine various ViTs with an EfficientNet [224], the latter being used for feature extraction. Due to its demonstrated strength in many tasks, the transformer architecture [259] is often employed to capture", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2091 }, { "text": "temporal incoherence. For instance, in a number of studies, the transformer is used together with 3D CNNs to extract temporal features [35], [36], [260]. Choi et al. [36] propose a complex framework that utilizes latent features from a pre- trained StyleGAN model [261], which are further encoded with GRUs. Cai et al. [34] employ the masked autoencoder", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2092 }, { "text": "pre-training framework to learn facial representations by guiding the masking strategy to focus on hiding face in- formation. The encoder is further fine-tuned on deepfake detection. Another interesting direction is to formulate the prob- lem as a graph classification task. Tan et al. [262] extract", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2093 }, { "text": "embeddings associated with the actions of different facial elements, systematically arrange them in a graph, and then employ a GCN to classify the video. Xu et al. [263] propose a novel strategy: to randomly sample frames from a video and combine them into a single image, called thumbnail. Then, the thumbnail is processed by a Swin Transformer [264] to", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2094 }, { "text": "obtain feature embeddings. Finally, these are fed into a GCN to capture any inconsistency and thus identify fake videos. 5.3 Audio 5.3.1 CNN-based methods The ability of CNNs to extract local features allows them to achieve competitive results in spoofed audio detection. Tak et al. [265] bring small modifications, such as fixed sinc", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2095 }, { "text": "filters, to the RawNet2 architecture and use it for spoofed speech detection. The same base architecture is further im- proved by Wang et al. [266] with orthogonal convolutions and temporal convolution networks (TCNs) to enhance the discrimination capability. In contrast, Conti et al. [267] intro- duce a new pipeline architecture that uses a Speech Emotion", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2096 }, { "text": "Recognition (SER) system as the feature extractor, and a Random Forest as the final classifier. Emotion features are extracted from an intermediate layer of a 3D-Convolutional Recurrent Neural Network. Hua et al. [268] introduce the Time-domain Synthetic Speech Detection Net (TSSDNet), an end-to-end framework that considers Inception-style con-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2097 }, { "text": "volutions and ResNet-like skip connections. The resulting architectures obtain state-of-the-art results on the ASVspoof 2019 dataset. 5.3.2 GNN-based methods Some models use the ability of GNNs to model relationships between entities in order to enhance spoofed speech detec- tion. Graph attention networks (GATs) are used in [269]\u2013", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2098 }, { "text": "and spectral sub-graphs, while Jung et al. [271] propose to combine the two sub-graphs into a single heterogeneous graph via a heterogeneous attention mechanism. Chen et al. [272] use a GCN to model the relationships from a graph constructed from patches of a spectrogram, outperforming competing models.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2099 }, { "text": "5.3.3 Transformer-based methods A growing number of methods use transformers for the synthesized speech detection task. Bartusiak et al. [273] em- ploy a compact convolutional transformer (CCT) to extract feature maps with a convolutional block from supplied spectrograms and further analyze them, after concatenation,", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2100 }, { "text": "using an attention mechanism. The CCT is extended in [274] to produce a compact attribution transformer (CAT) for the speech synthesizer attribution task, which aims to identify the tool/method that was used to synthesize the speech input. The proposed method also uses spectrograms as input, further producing a probability distribution over", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2101 }, { "text": "the set of known synthesizers. Mart\u00b4\u0131n-Do\u02dcnas et al. [275] present a model trained in a self-supervised manner, employing representations from different transformer layers of a pre-trained wav2vec 2.0 model to detect spoofed speech. The intermediate represen- tations from these layers are used to construct a vector for", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2102 }, { "text": "each time step. wav2vec 2.0 is also used by Cai et al. [276], who address partially fake audio detection. They identify fake audio segments by discovering the discontinuity be- tween them. Features are extracted with the aforementioned model, and frame embeddings are obtained by a ResNet- 1D module, while transformer-based encoders capture the", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2103 }, { "text": "context of the frames with respect to the sequence. Zhang et al. [277] aggregate a transformer architecture and a residual network, where the ability of the transformer to model long-term dependencies allows it to find corre- lations between audio frames. Different data augmentation techniques are used to increase the size of the training set", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2104 }, { "text": "and the final score is computed with a logistic regression meta-classifier that takes the scores from multiple models. Rawformer [278] aims to improve AASIST [271] by replac- ing the GAT with a transformer encoder. A positional aggre- gator augments the feature maps obtained by the RawNet2 feature extractor with positional information.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2105 }, { "text": "5.3.4 Other methods A method based on monitoring the behavior of neurons from a speaker recognition (SR) model is designed by Wang et al. [279]. The activated neurons from convolutional and fully-connected layers are used as feature vectors in the training process of a shallow network that classifies the", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2106 }, { "text": "IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. 50, NO. 1, NOVEMBER 2024 12 introduce Radian Weight Modification (RWM), a continual learning method that adjusts the direction of the gradient based on the means of the intra-class cosine distances of the samples from the current batch.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2107 }, { "text": "5.4 Multimodal 5.4.1 CNN-based methods Recent advances in deepfake and multimedia manipula- tion detection focus on combining audio-visual elements to improve model robustness and accuracy [25], [281], [282]. In Multimodaltrace [281], a ResNet-based framework blends audio and visual features, both independently and", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2108 }, { "text": "jointly, to enhance deepfake detection, offering insights into model focus areas using integrated gradient analysis. Similarly, Cozzolino et al. [25] propose a Person-of-Interest detector which leverages unique identity markers through contrastive learning with ResNet-50, excelling at detecting inconsistencies in low-quality videos. Kihal et al. [282] in-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2109 }, { "text": "troduce VTA-CNN-RF, a deep multimodal spam detection system, achieving over 98% precision in text, image, and audio spam classification using CNNs and Random Forests. 5.4.2 Transformer-based methods Recent deepfake detection frameworks based on attention mechanisms exploit both audio and visual cues to effec-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2110 }, { "text": "tively identify manipulated content [24], [283]\u2013[291]. Zhou et al. [24] propose a joint audio-visual detection method that leverages the synchronization between modalities, sig- nificantly boosting detection accuracy by late-fusing joint predictions with inter-attention mechanisms. Building on this approach, Oorloff et al. [283] develop a two-stage audio-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2111 }, { "text": "visual feature fusion method, using contrastive learning and autoencoders in the initial phase to capture audio-visual correspondences, followed by fine-tuning of transformer- based encoders for precise deepfake classification. Addi- tional frameworks that employ audio-visual cues have been proposed. For example, Salvi et al. [284] introduce a frame-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2112 }, { "text": "work analyzing audio-visual feature discrepancies over time, uniquely trained on separate monomodal datasets to identify unseen deepfakes. Similarly, AVFakeNet [285] is a unified model with dense Swin Transformer modules, which aptly handles variations in facial poses, lighting, and demographic diversity. Asha et al. [286] propose an", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2113 }, { "text": "ensemble-based D-Fence model, utilizing cross-modal atten- tion and self-attenuated neural networks to emphasize cor- relations between visual and audio elements for improved detection accuracy. For both intra and inter modality deepfake detection, Liu et al. [287] introduce the Forgery Clues Magnification", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2114 }, { "text": "Transformer (FCMT), which amplifies both intra-modal and cross-modal forgery cues through a distribution difference- based inconsistency computing module. Feng et al. [288] tackle audio-visual inconsistencies through an anomaly de- tection method that trains autoregressive transformers to flag low-probability sequences, using a joint ResNet-18 and", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2115 }, { "text": "VGG-M encoder. Zou et al. [289] advance cross-modality and within-modality regularization by aligning distinct audio and visual signals through multimodal transformers, while Nie et al. [290] introduce FRADE, which relies on adaptive forgery-aware injection and audio-distilled cross-modal in- teraction to effectively bridge the audio-visual domain gap.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2116 }, { "text": "TABLE 2 Datasets that are commonly used in deepfake detection literature, separated by domain. AV stands for audio-video (multimodal). Dataset #Real #Fake Resolution/frequency Image DFFD [292] 58,703 240,336 250\u00d7250 - 1024\u00d71024 FakeSpotter [293] 6,000 6,000 224\u00d7224 ForgeryNet [294] 1,438,201 1,457,861 240\u00d7240 - 1080\u00d71080", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2117 }, { "text": "DiffusionFace [295] 30,000 600,000 256\u00d7256 Video FaceForensics++ [296] 1,000 4,000 512\u00d7512 DeeperForensics [297] 48,475 11,000 1920\u00d71080 Celeb-DF [298] 590 5,639 256\u00d7256 WildDeepfake [299] 3,805 3,509 varying DeepFake-TIMIT [300] 0 620 64\u00d764/128\u00d7128 UADFV [301] 98 98 294\u00d7500 GenVideo [302] 1,224,511 1,089,671 224\u00d7224 - 1280\u00d72048", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2118 }, { "text": "Audio WaveFake [303] 0 117,985 16 kHz ASVspoof 2019-LA [304] 12,483 108,978 16 kHz ASVspoof 2021-LA [305] 16492 148148 16 kHz ASVspoof 2021-DF [305] 20,637 572,616 16 kHz In-the-Wild [306] 19,963 11,816 16 kHz ADD 2022 [307] 91,464 358,082 16 kHz ADD 2023 [308] 243,194 273,874 16 kHz FoR [309] 111,000", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2119 }, { "text": "87,285 16 kHz MLAAD [310] 0 154,000 22 kHz AV FakeAVCeleb [311] 500 19,500 224\u00d7224 LAV-DF [312] 36,431 99,873 224\u00d7224 DFDC [313] 23,654 104,500 1920\u00d71080/1080\u00d71920 Moreover, Yang et al. [314] introduce AVoiD-DF, a model based on a temporal-spatial encoder and a multimodal joint decoder. AVoiD-DF captures inter-modal and intra-modal", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2120 }, { "text": "disharmony, achieving good performance across various forgery techniques. 5.4.3 Other methods Hosler et al. [315] introduce a method for detecting deep- fakes by analyzing emotional consistency in human faces and voices using LSTM networks. By predicting emotions from audio and video features, the approach identifies un-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2121 }, { "text": "natural emotional patterns to flag deepfake media. 6 DATASETS AND REPORTED RESULTS In Table 2, we present the most frequently used datasets for deepfake detection, along with the number of real and fake samples, as well as the resolution (for visual datasets) or the bit rate (for audio datasets). We next de-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2122 }, { "text": "scribe the main steps that are usually employed to build deepfake datasets. The first step of creating a dataset for deepfake detection is collecting the real data. Except for Dolhansky et al. [313], who create the original data by recording movies of paid actors, the basic procedure is to scrape the Internet for videos, especially YouTube.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2123 }, { "text": "Even image datasets use frames extracted from videos. After acquiring real data, various deepfake methods are applied to generate the fake samples. Due to their ex- cellent trade-off between performance and speed, GANs are adopted for the creation of most datasets, e.g. Celeb- DF [298], DeepFake-TIMIT [300], DFFD [292], FakeSpot-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2124 }, { "text": "ter [293] and FakeAVCeleb [311]. VAEs represent the method of choice only for a few datasets, e.g. DeeperForen- sics [297] and ASVspoof 2019-LA [304]. Diffusion models are recent and powerful generative methods, yet they re- quire more computation. Hence, only a couple of recent datasets employ them to create the fake samples, e.g. Gen-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2125 }, { "text": "IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. 50, NO. 1, NOVEMBER 2024 13 TABLE 3 Results of top scoring image deepfake detection methods on DFFD [292], DiffusionFace [295], ForgeryNet [294]. Dataset Method Accuracy AUC DFFD [292] BNext-M [221] 99.18% 0.9994 VGG-16 [292] - 0.9967", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2126 }, { "text": "XceptionNet [292] - 0.9964 DiffusionFace [295] GramNet [227] 62.60% 0.7250 GFF [228] 61.10% 0.7250 RECCE [233] 64.40% 0.7130 F3Net [229] 59.80% 0.6960 ForgeryNet [294] SNRFilters-Xception [240] 81.09% 0.9052 GramNet [227] 80.89% 0.9020 F3Net [229] 80.86% 0.9015 XceptionNet [296] 80.78% 0.9012 TABLE 4", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2127 }, { "text": "Results of top scoring video deepfake detection methods on FaceForensics++ [296], DFDC [313] and Celeb-DF [298] datasets. Dataset Method Accuracy AUC FaceForensics++ [296] TALL++ [263] 98.65% 0.9987 LipForensics [246] 98.90% 0.9970 FADE [262] 92.89% 0.9952 M2TR [32] 97.93% 0.9951 App.+Beh. [241] 98.90%", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2128 }, { "text": "0.9900 RealForensics [247] - 0.9900 DFDC [313] Efficient ViT [258] - 0.9510 TALL++ [263] - 0.9068 CNN Ensemble [255] - 0.8782 RealForensics [247] - 0.7590 LipForensics [246] - 0.7350 Celeb-DF [298] App.+Beh. [241] 98.50% 0.9900 FInfer [250] 90.47% 0.9330 RealForensics [247] - 0.8690 LipForensics [246]", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2129 }, { "text": "- 0.8240 method, e.g. GenVideo [302], ForgeryNet [294], FaceForen- sics++ [296], DFDC [313], LAV-DF [312], WaveFake [303], ASVspoof 2019-LA [304], ASVspoof 2021-LA/DF [305], FoR [309] and MLAAD [310]. A few visual datasets [296], [302] rely on online tools to create deepfakes, the most popular tool being FaceSwap6.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2130 }, { "text": "Depending on the input modality, different metrics are commonly reported. For the visual modalities, the most fre- quent metric is the area under the curve (AUC). Given that deepfake detection is a binary classification task, the AUC score can illustrate the ability of the model to differentiate between real and fake samples. The AUC is obtained by", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2131 }, { "text": "plotting the True Positive Rate against the False Positive Rate for multiple thresholds, and then computing the area under the resulting curve. Accuracy is an alternative metric that can be used to assess the overall performance of a deepfake detection model. Nevertheless, deepfake detection datasets are usually imbalanced, making accuracy a less pre-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2132 }, { "text": "ferred option. For the audio modality, models are regularly evaluated via the equal error rate (EER). Its popularity is given by the robustness to class imbalance, while equally assessing false positives and false negatives. EER is com- puted by finding the intersection of the False Acceptance Rate and the False Rejection Rate.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2133 }, { "text": "6.1 Results on Popular Benchmarks Image. Table 3 includes top accuracy and AUC scores on three commonly-used datasets of deepfake images, namely DFFD [292], DiffusionFace [295] and ForgeryNet [294]. The 6. https://github.com/deepfakes/faceswap TABLE 5 Results of top scoring audio deepfake detection methods on the", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2134 }, { "text": "ASVspoof 2019-LA [304] and ASVspoof 2021-LA [305] datasets. Dataset Method EER (%) ASVspoof 2019-LA [304] GCN [272] 0.58 Rawformer [278] 0.59 AASIST [271] 0.83 RawGAT [270] 1.06 TO-RawNet [266] 1.58 TSSDNet [268] 1.64 ASVspoof 2021-LA [305] wav2vec2+AASIST [316] 0.82 wav2vec2+MLP [275] 3.54 TO-RawNet [266]", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2135 }, { "text": "3.70 Rawformer [278] 4.53 AASIST [271] 9.15 TABLE 6 Results of top scoring multimodal deepfake detection methods on the FakeAVCeleb [311] and DFDC [313] datasets. Dataset Method Accuracy AUC FakeAVCeleb [311] FRADE [290] 98.60% 0.9980 AVFF [283] 98.60% 0.9910 MIS-AVoiDD [317] 96.20% 0.9730 PVASS-MDD [318]", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2136 }, { "text": "95.70% 0.9730 SSVF [288] 94.20% 0.9450 MRDF [289] 94.05% 0.9243 DFDC [313] FRADE [290] 97.20% 0.9900 PVASS-MDD [318] 96.30% 0.9890 AVoiD-DF [314] 91.40% 0.9480 AVA-CL [291] 84.20% 0.8864 AVFakeNet [285] 82.80% 0.8620 results suggest that the oldest dataset, DFFD, has become saturated due to advancements in recent deepfake detectors.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2137 }, { "text": "In contrast, the most recent dataset, DiffusionFace, featuring faces generated by diffusion models, poses a significantly greater challenge for state-of-the-art detectors. This high- lights the need for future developments in deepfake detec- tion to effectively differentiate between genuine faces and", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2138 }, { "text": "those synthesized by diffusion models. Video. Table 4 provides the performance levels of a few of the most effective methods for video deepfake detection on three distinctive datasets: FaceForensics++ [296], DeepFake Detection Challenge (DFDC) [313] and Celeb-DF [298]. The main metric in this area is the AUC, but we also report the", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2139 }, { "text": "accuracy, whenever it is available. All the included models are outstanding, most of them almost achieving flawless performance. This is not only true for the newest methods that employ the most recent trends (such as transformers), but also for the preceding ones, that solely utilize CNNs. Nevertheless, on the more difficult datasets (DFDC and", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2140 }, { "text": "Celeb-DF), it can be observed that ViT-based architectures are superior. Audio. In Table 5, we report the Equal Error Rate (EER) val- ues of top audio deepfake detection methods on ASVspoof 2019-LA [304] and ASVspoof 2021-LA [305], two of the most popular audio deepfake detection datasets. GNN-based", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2141 }, { "text": "methods achieve the lowest EER values on both datasets, demonstrating their effectiveness in the synthesized speech detection task. Multimodal. In Table 6, we present the performance levels of top deepfake detection methods on two popular mul- timodal datasets, namely FakeAVCeleb [311] and DFDC [313]. For each method, we report the performance in terms", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2142 }, { "text": "IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. 50, NO. 1, NOVEMBER 2024 14 Fig. 3. Randomly sampled frames captured from the fake videos in- cluded in BioDeepAV exhibit a high level of realism. Best viewed in color. with FRADE [290] and AVFF [283] achieving the highest accuracy rates, both at 98.60%. On the DFDC dataset,", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2143 }, { "text": "we include five methods, with FRADE and PVASS-MDD demonstrating top performance, attaining accuracy rates of 97.20% and 96.30%, respectively. The reported results hint towards important advancements in multimodal methods, with recent methods nearly saturating the benchmarks. 6.2 Proposed Benchmark We create a new dataset, called BioDeepAV7, to assess", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2144 }, { "text": "the out-of-domain generalization capabilities of deepfake detection models. Our primary focus is on generating videos featuring talking faces, but we also include audio-video examples with audio-only manipulations. Figure 3 depicts a few frames from various deepfake videos in BioDeepAV. Generated Data. We generate over 1,600 deepfake videos", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2145 }, { "text": "using four recent methods specialized in talking-face syn- thesis [181], [184], [319], [320]. These approaches base their solutions on the recent development of diffusion mod- els [181], [319], NeRFs [184] and Gaussian Splatting [320]. We use three face image sources to sample target identities. First, we create 300 synthetic faces using RealVisXLv58 and", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2146 }, { "text": "supplement these with faces from the LAION-Face [321] and HDTF [322] datasets. Three of the methods [181], [184], [320] also require head motion information as a condition- ing signal, which we obtain from the HDTF [322] dataset. In addition to face images and motion cues, all methods also require an audio file to condition their talking-face", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2147 }, { "text": "generation. For this, we use the samples from a dataset of English dialects [323], the audio from the HDTF [322] dataset, and over 700 deepfake audio samples created by us. To generate these synthetic audio samples, we employ StyleTTS [324], SSR-Speech [325] and YourTTS [158], which support both text-to-speech synthesis [158], [324], [325] and", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2148 }, { "text": "voice conversion [158]. We source text prompts for text-to- speech synthesis from the LibriTTS dataset [326], and use the speakers from this dataset for voice conversion, with tar- get audio sourced from the dataset of English dialects [323]. Real Data. We sample real videos for our experiments from", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2149 }, { "text": "two datasets, HDTF [322] and TalkingHead-1KH [327]. We include all available videos from HDTF, while sampling an additional 2,000 videos from TalkingHead-1KH. 7. Available at: https://github.com/CroitoruAlin/biodeep 8. https://civitai.com/models/139562/realvisxl-v50 TABLE 7 Results (in terms of AUC) of four state-of-the-art deepfake detectors on", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2150 }, { "text": "the original test sets versus BioDeepAV. UCF [200], RECCE [233], TALL [263], F3Net [229], StA [257] and XceptionNet [296] are originally tested on FaceForensics++ [296], while MRDF is originally tested on FakeAVCeleb [311]. Method Venue Original Test BioDeepAV StA [257] ArXiv 2024 0.9420 0.6195 XceptionNet [296]", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2151 }, { "text": "ICCV 2019 0.9637 0.5677 F3Net [229] ECCV 2020 0.9449 0.5010 RECCE [233] CVPR 2022 0.9422 0.5001 TALL [263] ICCV 2023 0.9987 0.4935 UCF [200] ICCV 2023 0.9527 0.4882 MRDF [289] ICASSP 2024 0.9243 0.5852 Experiments. We run the experiments using the Deep- fakeBench benchmark [328], which implements state-of-the-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2152 }, { "text": "art deepfake detectors. For our analysis, we choose three image-based detectors, namely UCF [200], RECCE [233] and a model based on XceptionNet [296], one detector applied on the frequency domain, namely F3Net [229], two video- based detectors, namely TALL [263] and StA [257], and one audio-visual detector, namely MRDF [289]. MRDF is not im-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2153 }, { "text": "plemented in the DeepFakeBench benchmark, so we employ the official implementation in our experiments. UCF [200], RECCE [233], TALL [263], F3Net [229], StA [257] and Xcep- tionNet [296] are trained on FaceForensics++ [296], while MRDF [289] is trained on FakeAVCeleb [311]. In Table 7, we report the video AUC of these detectors on BioDeepAV and", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2154 }, { "text": "their original test sets, respectively. The considered methods always surpass the 90% threshold when tested in-domain, yet all methods register drastic performance drops (higher than 30%) on BioDeepAV. The findings clearly demonstrate that current detectors struggle to identify the authenticity of talking faces generated by the novel (unseen) generative", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2155 }, { "text": "models included in BioDeepAV, highlighting the need for further research in this area. 7 CONCLUSIONS AND FUTURE DIRECTIONS In this paper, we reviewed deepfake generation and de- tection methods, constructing a comprehensive taxonomy of methods across image, video, audio and multimodal domains. After discussing the methods included in our", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2156 }, { "text": "taxonomy, we turned our attention to datasets used for deepfake detection, with a particular focus on the results reported by top performing models. Moreover, we evalu- ated some of the best methods on our novel benchmark, BioDeepAV, aiming to assess the generalization capacity of current deepfake detectors to out-of-distribution data. The", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2157 }, { "text": "results show that the distribution gap can greatly affect state-of-the-art deepfake detectors, pinpointing the need for more robust models. Future directions. Based on the observed gaps in deepfake literature, there are several directions which we recommend exploring in future work. The most important future di-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2158 }, { "text": "rection is the development of deepfake detectors that can generalize across multiple generative tools. Our new bench- mark, BioDeepAV, will come in handy to test the general- ization capacity of deepfake detection models in the future. Another area that is not sufficiently explored is the devel- opment of explainable deepfake detectors. Knowing when", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2159 }, { "text": "IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. 50, NO. 1, NOVEMBER 2024 15 often been disregarded. While current detectors mostly rely on deep neural networks, an important downside of such models is that they are unable to quantify their uncertainty. To this end, studying approaches to calibrate deepfake de-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2160 }, { "text": "tectors could lead to the development of enhanced models able to indicate when their prediction is uncertain. ACKNOWLEDGMENTS This work was supported by a grant of the Ministry of Research, Innovation and Digitization, CCCDI - UEFIS- CDI, project number PN-IV-P6-6.3-SOL-2024-2-0227, within PNCDI IV.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2161 }, { "text": "REFERENCES [1] F.-A. Croitoru, V. Hondru, R. T. Ionescu, et al., \u201cDiffusion Models in Vision: A Survey,\u201d IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023. [2] J. Chen, J. Yu, C. Gu, et al., \u201cPixart-\u03b1: Fast training of diffusion transformer for photorealistic text-to-image synthesis,\u201d in ICLR,", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2162 }, { "text": "2024. [3] R. Rombach, A. Blattmann, D. Lorenz, et al., \u201cHigh-Resolution Image Synthesis with Latent Diffusion Models,\u201d in CVPR, 2022. [4] D. Podell, Z. English, K. Lacey, et al., \u201cSDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis,\u201d in ICLR, 2024. [5] A. Sauer, K. Schwarz, and A. Geiger, \u201cStyleGAN-XL: Scaling", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2163 }, { "text": "StyleGAN to Large Diverse Datasets,\u201d in SIGGRAPH, 2022. [6] P. Esser, R. Rombach, and B. Ommer, \u201cTaming transformers for high-resolution image synthesis,\u201d in CVPR, 2021. [7] Y. Li, C. Ma, Y. Yan, et al., \u201c3D-Aware Face Swapping,\u201d in CVPR, 2023. [8] Z. Liu, M. Li, Y. Zhang, et al., \u201cFine-Grained Face Swapping Via", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2164 }, { "text": "Regional GAN Inversion,\u201d in CVPR, 2023. [9] J. Kim, J. Lee, and B. Zhang, \u201cSmooth-Swap: A Simple Enhance- ment for Face-Swapping with Smoothness,\u201d in CVPR, 2022. [10] K. Shiohara, X. Yang, and T. Taketomi, \u201cBlendFace: Re-designing Identity Encoders for Face-Swapping,\u201d in ICCV, 2023. [11] F. Paraperas Papantoniou, A. Lattas, et al., \u201cArc2Face: A Founda-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2165 }, { "text": "tion Model for ID-Consistent Human Faces,\u201d in ECCV, 2024. [12] W. Zhao, Y. Rao, W. Shi, et al., \u201cDiffSwap: High-Fidelity and Controllable Face Swapping via 3D-Aware Masked Diffusion,\u201d in CVPR, 2023. [13] R. Liu, B. Ma, W. Zhang, et al., \u201cTowards a simultaneous and granular identity-expression control in personalized face genera-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2166 }, { "text": "tion,\u201d in CVPR, 2024. [14] Y. Nirkin, Y. Keller, and T. Hassner, \u201cFSGAN: Subject Agnostic Face Swapping and Reenactment,\u201d in ICCV, 2019. [15] S. Bounareli, C. Tzelepis, V. Argyriou, et al., \u201cHyperReenact: One- Shot Reenactment via Jointly Learning to Refine and Retarget Faces,\u201d in ICCV, 2023. [16] T. Oorloff and Y. Yacoob, \u201cRobust One-Shot Face Video Re-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2167 }, { "text": "enactment using Hybrid Latent Spaces of StyleGAN2,\u201d in ICCV, 2023. [17] Y. Ma, S. Zhang, J. Wang, et al., \u201cDreamTalk: When Expressive Talking Head Generation Meets Diffusion Probabilistic Models,\u201d arXiv:2312.09767, 2023. [18] M. Xu, H. Li, Q. Su, et al., \u201cHallo: Hierarchical audio-driven visual synthesis for portrait image animation,\u201d arXiv:2406.08801, 2024.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2168 }, { "text": "bling block to improving deepfake detection generalization,\u201d in CVPR, 2023. [22] L. Lin, X. He, Y. Ju, et al., \u201cPreserving fairness generalization in deepfake detection,\u201d in CVPR, 2024. [23] Y. Xu, J. Liang, L. Sheng, et al., \u201cLearning spatiotemporal in- consistency via thumbnail layout for face deepfake detection,\u201d", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2169 }, { "text": "International Journal of Computer Vision, 2024. [24] Y. Zhou and S.-N. Lim, \u201cJoint audio-visual deepfake detection,\u201d in ICCV, 2021. [25] D. Cozzolino, A. Pianese, M. Nie\u00dfner, et al., \u201cAudio-visual person-of-interest deepfake detection,\u201d in CVPR, 2023. [26] B. Goyal, N. S. Gill, P. Gulia, et al., \u201cDetection of fake accounts on", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2170 }, { "text": "social media using multimodal data with deep learning,\u201d IEEE Transactions on Computational Social Systems, 2023. [27] R. Natsume, T. Yatagawa, and S. Morishima, \u201cRSGAN: Face swapping and editing using face and hair representation in latent spaces,\u201d in SIGGRAPH, 2018. [28] I. Skorokhodov, S. Tulyakov, and M. Elhoseiny, \u201cStyleGAN-V: A", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2171 }, { "text": "Continuous Video Generator with the Price, Image Quality and Perks of StyleGAN2,\u201d in CVPR, 2022. [29] C. Kim, J. Lee, S. Joung, et al., \u201cInstantFamily: Masked Atten- tion for Zero-shot Multi-ID Image Generation,\u201d arXiv:2404.19427, 2024. [30] Y. Wang, W. Zhang, J. Zheng, et al., \u201cHigh-fidelity person-centric", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2172 }, { "text": "subject-to-image synthesis,\u201d in CVPR, 2024. [31] M. Stypu\u0142kowski, K. Vougioukas, S. He, et al., \u201cDiffused Heads: Diffusion Models Beat GANs on Talking-Face Generation,\u201d in WACV, 2024. [32] J. Wang, Z. Wu, W. Ouyang, et al., \u201cM2TR: Multi-modal Multi- scale Transformers for Deepfake Detection,\u201d in ICMR, 2022.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2173 }, { "text": "for more general video face forgery detection,\u201d in ICCV, 2021. [36] J. Choi, T. Kim, Y. Jeong, et al., \u201cExploiting Style Latent Flows for Generalizing Deepfake Video Detection,\u201d in CVPR, 2024. [37] R. Shao, T. Wu, and Z. Liu, \u201cDetecting and recovering sequential deepfake manipulation,\u201d in ECCV, 2022.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2174 }, { "text": "detection techniques using deep learning,\u201d in ICNGIS, 2022. [41] A. Heidari, N. J. Navimipour, H. Dag, et al., \u201cDeepfake detection using deep learning methods: A systematic and comprehensive review,\u201d WIREs Data Mining and Knowledge Discovery, 2024. [42] A. Kaur, A. Noori Hoshyar, V. Saikrishna, et al., \u201cDeepfake video", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2175 }, { "text": "detection: challenges and opportunities,\u201d Artificial Intelligence Review, 2024. [43] W. Lei, J. Wang, F. Ma, et al., \u201cA comprehensive survey on human video generation: Challenges, methods, and insights,\u201d arXiv:2407.08428, 2024. [44] M. Li, Y. Ahmadiadli, and X.-P. Zhang, \u201cAudio anti-spoofing detection: A survey,\u201d arXiv:2404.13914, 2024.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2176 }, { "text": "of deepfake: Generation, detection, datasets, and opportunities,\u201d Neurocomputing, 2022. [50] J. Yi, C. Wang, J. Tao, et al., \u201cAudio deepfake detection: A survey,\u201d arXiv:2308.14970, 2023. [51] T. Zhang, \u201cDeepfake generation and detection, a survey,\u201d Multi- media Tools and Applications, 2022. [52] T. Brooks, B. Peebles, C. Holmes, et al., \u201cVideo generation models", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2177 }, { "text": "as world simulators,\u201d tech. rep., OpenAI, 2024. [53] A. Q. Nichol, P. Dhariwal, A. Ramesh, et al., \u201cGLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models,\u201d in ICML, 2022. [54] Y. Shen, P. Luo, J. Yan, et al., \u201cFaceID-GAN: Learning a Symmetry Three-Player GAN for Identity-Preserving Face Synthesis,\u201d in", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2178 }, { "text": "CVPR, 2018. [55] T. Karras, T. Aila, S. Laine, et al., \u201cProgressive Growing of GANs for Improved Quality, Stability, and Variation,\u201d in ICLR, 2018. [56] T. Karras, S. Laine, and T. Aila, \u201cA style-based generator archi- tecture for generative adversarial networks,\u201d IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2179 }, { "text": "IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. 50, NO. 1, NOVEMBER 2024 16 [57] T. Karras, S. Laine, M. Aittala, et al., \u201cAnalyzing and Improving the Image Quality of StyleGAN,\u201d in CVPR, 2020. [58] T. Karras, M. Aittala, S. Laine, et al., \u201cAlias-free generative adver- sarial networks,\u201d in NeurIPS, 2021.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2180 }, { "text": "with adaptive instance normalization,\u201d in ICCV, 2017. [62] A. Sauer, K. Chitta, J. M\u00a8uller, et al., \u201cProjected GANs converge faster,\u201d in NeurIPS, 2021. [63] P. Dhariwal and A. Nichol, \u201cDiffusion models beat GANs on image synthesis,\u201d in NeurIPS, 2021. [64] I. Goodfellow, J. Pouget-Abadie, M. Mirza, et al., \u201cGenerative", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2181 }, { "text": "adversarial nets,\u201d in NeurIPS, 2014. [65] J. Bao, D. Chen, F. Wen, et al., \u201cTowards open-set identity pre- serving face synthesis,\u201d in CVPR, 2018. [66] R. Chen, X. Chen, B. Ni, et al., \u201cSimSwap: An Efficient Framework For High Fidelity Face Swapping,\u201d in ACMMM, 2020. [67] X. Ren, X. Chen, P. Yao, et al., \u201cReinforced disentanglement for", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2182 }, { "text": "face swapping without skip connection,\u201d in ICCV, 2023. [68] F. Rosberg, E. Aksoy, F. Alonso-Fernandez, et al., \u201cFaceDancer: Pose- and Occlusion-Aware High Fidelity Face Swapping,\u201d in WACV, 2023. [69] L. Li, J. Bao, H. Yang, et al., \u201cAdvancing high fidelity identity swapping for forgery detection,\u201d in CVPR, 2020.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2183 }, { "text": "Fidelity and Accurate Face Swapping,\u201d in CVPRW, 2023. [73] G. Yuan, M. Li, Y. Zhang, et al., \u201cReliableSwap: Boosting General Face Swapping Via Reliable Supervision,\u201d arXiv:2306.05356, 2023. [74] W. Cao, T. Wang, A. Dong, et al., \u201cTransFS: Face Swapping Using Transformer,\u201d in FG, 2023. [75] K. He, X. Chen, S. Xie, et al., \u201cMasked autoencoders are scalable", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2184 }, { "text": "vision learners,\u201d in CVPR, 2022. [76] G. Gao, H. Huang, C. Fu, et al., \u201cInformation bottleneck disentan- glement for identity swapping,\u201d in CVPR, 2021. [77] Y. Zhu, Q. Li, J. Wang, et al., \u201cOne shot face swapping on megapixels,\u201d in CVPR, 2021. [78] D. Jiang, D. Song, R. Tong, et al., \u201cStyleIPSB: Identity-Preserving", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2185 }, { "text": "Semantic Basis of StyleGAN for High Fidelity Face Swapping,\u201d in CVPR, 2023. [79] E. R. Chan, M. Monteiro, P. Kellnhofer, et al., \u201cpi-GAN: Periodic Implicit Generative Adversarial Networks for 3D-Aware Image Synthesis,\u201d in CVPR, 2021. [80] J. R. A. Moniz, C. Beckham, S. Rajotte, et al., \u201cUnsupervised", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2186 }, { "text": "Depth Estimation, 3D Face Rotation and Replacement,\u201d in NeurIPS, 2018. [81] Q. Sun, A. Tewari, W. Xu, et al., \u201cA hybrid model for identity obfuscation by face replacement,\u201d in ECCV, 2018. [82] Z. Chen, L. Xie, S. Pang, et al., \u201cMagDR: Mask-guided Detection and Reconstruction for Defending Deepfakes,\u201d in CVPR, 2021.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2187 }, { "text": "Y. Choi, Y. Uh, J. Yoo, et al., \u201cStarGAN v2: Diverse Image Synthe- sis for Multiple Domains,\u201d in CVPR, 2020. [86] G.-S. Hsu, C.-H. Tsai, and H.-Y. Wu, \u201cDual-generator face reen- actment,\u201d in CVPR, 2022. [87] O. Tov, Y. Alaluf, Y. Nitzan, et al., \u201cDesigning an encoder for StyleGAN image manipulation,\u201d ACM Transactions on Graphics,", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2188 }, { "text": "2021. [88] A. Suwa\u0142a, B. W\u00b4ojcik, M. Proszewska, et al., \u201cFace Identity-Aware Disentanglement in StyleGAN,\u201d in WACV, 2024. [89] C. Saharia, W. Chan, S. Saxena, et al., \u201cPhotorealistic Text-to- Image Diffusion Models with Deep Language Understanding,\u201d in NeurIPS, 2022. [90] N. Ruiz, Y. Li, V. Jampani, et al., \u201cDreamBooth: Fine Tuning Text-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2189 }, { "text": "to-Image Diffusion Models for Subject-Driven Generation,\u201d in CVPR, 2023. [91] Z. Chen, S. Fang, W. Liu, et al., \u201cDreamIdentity: Enhanced Ed- itability for Efficient Face-Identity Preserved Image Generation,\u201d in AAAI, 2024. [92] X. Peng, J. Zhu, B. Jiang, et al., \u201cPortraitBooth: A Versatile Portrait", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2190 }, { "text": "Model for Fast Identity-Preserved Personalization,\u201d in CVPR, 2024. [93] J. Ma, J. Liang, C. Chen, et al., \u201cSubject-Diffusion: Open Domain Personalized Text-to-Image Generation without Test-time Fine- tuning,\u201d in SIGGRAPH, 2024. [94] Y. Liu, C. Yu, L. Shang, et al., \u201cFaceChain: A Playground for Human-centric Artificial Intelligence Generated Content,\u201d", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2191 }, { "text": "arXiv:2308.14256, 2023. [95] F. Boutros, J. H. Grebe, A. Kuijper, et al., \u201cIDiff-Face: Synthetic- based Face Recognition through Fizzy Identity-Conditioned Dif- fusion Model,\u201d in ICCV, 2023. [96] Y. Han, J. Zhang, J. Zhu, et al., \u201cA Generalist FaceX via Learning Unified Facial Representation,\u201d arXiv:2401.00551, 2023.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2192 }, { "text": "Personalization via ID-semantics Decoupling Paradigm,\u201d arXiv:2403.11781, 2024. [100] Q. Wang, X. Jia, X. Li, et al., \u201cStableIdentity: Inserting Anybody into Anywhere at First Sight,\u201d arXiv:2401.15975, 2024. [101] Z. Guo, Y. Wu, Z. Chen, et al., \u201cPuLID: Pure and Lightning ID Customization via Contrastive Alignment,\u201d in NeurIPS, 2024.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2193 }, { "text": "and Face Identity for Personalized Text-to-Image Generation,\u201d in ECCV, 2024. [111] R. Gal, Y. Alaluf, Y. Atzmon, et al., \u201cAn image is worth one word: Personalizing text-to-image generation using textual inversion,\u201d in ICLR, 2023. [112] N. Ruiz, Y. Li, V. Jampani, et al., \u201cHyperDreamBooth: Hyper- Networks for Fast Personalization of Text-to-Image Models,\u201d in", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2194 }, { "text": "CVPR, 2024. [113] W. Chen, H. Hu, Y. Li, et al., \u201cSubject-driven text-to-image gener- ation via apprenticeship learning,\u201d in NeurIPS, 2023. [114] E. J. Hu, Y. Shen, P. Wallis, et al., \u201cLoRA: Low-Rank Adaptation of Large Language Models,\u201d in ICLR, 2022. [115] D. Bitouk, N. Kumar, S. Dhillon, et al., \u201cFace swapping: automat-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2195 }, { "text": "ically replacing faces in photographs,\u201d in SIGGRAPH, 2008. [116] I. Korshunova, W. Shi, J. Dambre, et al., \u201cFast face-swap using convolutional neural networks,\u201d in ICCV, 2017. [117] Y. Wang, X. Chen, J. Zhu, et al., \u201cHifiFace: 3D Shape and Semantic Prior Guided High Fidelity Face Swapping,\u201d in IJCAI, 2021.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2196 }, { "text": "IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. 50, NO. 1, NOVEMBER 2024 17 [119] X. Li, X. Hou, and C. C. Loy, \u201cWhen StyleGAN Meets Stable Diffusion: a W+ Adapter for Personalized Image Generation,\u201d in CVPR, 2024. [120] Z. Xu, H. Zhou, Z. Hong, et al., \u201cStyleSwap: Style-Based Genera-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2197 }, { "text": "tor Empowers Robust Face Swapping,\u201d in ECCV, 2022. [121] Y. Gao, Y. Zhou, J. Wang, et al., \u201cHigh-Fidelity and Freely Con- trollable Talking Head Video Generation,\u201d in CVPR, 2023. [122] Y. Tian, J. Ren, M. Chai, et al., \u201cA Good Image Generator Is What You Need for High-Resolution Video Synthesis,\u201d in ICLR, 2021.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2198 }, { "text": "sis: Disentangled Controls for StyleGAN Image Generation,\u201d in CVPR, 2021. [126] S. Yu, J. Tack, S. Mo, et al., \u201cGenerating Videos with Dynamics- aware Implicit Generative Adversarial Networks,\u201d in ICLR, 2022. [127] V. Sitzmann, J. Martel, A. Bergman, et al., \u201cImplicit Neural Repre- sentations with Periodic Activation Functions,\u201d in NeurIPS, 2020.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2199 }, { "text": "in CVPR, 2023. [131] A. Blattmann, T. Dockhorn, S. Kulal, et al., \u201cStable Video Diffu- sion: Scaling Latent Video Diffusion Models to Large Datasets,\u201d arXiv:2311.15127, 2023. [132] J. Z. Wu, Y. Ge, X. Wang, et al., \u201cTune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video Generation,\u201d in", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2200 }, { "text": "ICCV, 2023. [133] F. Bao, C. Xiang, G. Yue, et al., \u201cVidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models,\u201d arXiv:2405.04233, 2024. [134] Y. Guo, C. Yang, A. Rao, et al., \u201cAnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning,\u201d in ICLR, 2024.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2201 }, { "text": "Video Generation without Text-Video Data,\u201d in ICLR, 2022. [138] A. Ramesh, P. Dhariwal, A. Nichol, et al., \u201cHierarchi- cal Text-Conditional Image Generation with CLIP Latents,\u201d arXiv:2204.06125, 2022. [139] J. Ho, W. Chan, C. Saharia, et al., \u201cImagen Video: High Defini- tion Video Generation with Diffusion Models,\u201d arXiv:2210.02303,", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2202 }, { "text": "2022. [140] S. Bounareli, C. Tzelepis, V. Argyriou, et al., \u201cDiffusionAct: Con- trollable Diffusion Autoencoder for One-shot Face Reenactment,\u201d arXiv:2403.17217, 2024. [141] Y. Ma, H. Liu, H. Wang, et al., \u201cFollow-Your-Emoji: Fine- Controllable and Expressive Freestyle Portrait Animation,\u201d arXiv:2406.01900, 2024.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2203 }, { "text": "Representation Transformer for Face Reenactment from Factor- ized Appearance Head-pose and Facial Expression Features,\u201d in CVPR, 2024. [145] R. Villegas, M. Babaeizadeh, P.-J. Kindermans, et al., \u201cPhenaki: Variable Length Video Generation From Open Domain Textual Description,\u201d in ICLR, 2023. [146] A. Van Den Oord, O. Vinyals, and K. Kavukcuoglu, \u201cNeural", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2204 }, { "text": "discrete representation learning,\u201d in NeurIPS, 2017. [147] L. Yu, Y. Cheng, K. Sohn, et al., \u201cMAGVIT: Masked Generative Video Transformer,\u201d in CVPR, 2023. [148] W. Hong, M. Ding, W. Zheng, et al., \u201cCogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers,\u201d in ICLR, 2023. [149] Y. Jiang, S. Yang, T. L. Koh, et al., \u201cText2Performer: Text-Driven", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2205 }, { "text": "Human Video Generation,\u201d in ICCV, 2023. [150] W. Yan, Y. Zhang, P. Abbeel, et al., \u201cVideoGPT: Video Generation using VQ-VAE and Transformers,\u201d arXiv:2104.10157, 2021. [151] B. Mildenhall, P. P. Srinivasan, M. Tancik, et al., \u201cNeRF: Rep- resenting scenes as neural radiance fields for view synthesis,\u201d", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2206 }, { "text": "Communications of the ACM, 2021. [152] J. Thies, M. Zollh\u00a8ofer, and M. Nie\u00dfner, \u201cDeferred Neural Render- ing: Image Synthesis using Neural Textures,\u201d ACM Transactions on Graphics, 2019. [153] J. Thies, M. Zollhofer, M. Stamminger, et al., \u201cFace2Face: Real- time Face Capture and Reenactment of RGB Videos,\u201d in CVPR,", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2207 }, { "text": "2016. [154] K. Yang, K. Chen, D. Guo, et al., \u201cFace2Face \u03c1: Real-Time High- Resolution One-Shot Face Reenactment,\u201d in ECCV, 2022. [155] B. Zhang, C. Qi, P. Zhang, et al., \u201cMetaPortrait: Identity- Preserving Talking Head Generation with Fast Personalized Adaptation,\u201d in CVPR, 2023. [156] W. Li, L. Zhang, D. Wang, et al., \u201cOne-Shot High-Fidelity Talking-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2208 }, { "text": "Head Synthesis with Deformable Neural Radiance Field,\u201d in CVPR, 2023. [157] J. Kim, J. Kong, and J. Son, \u201cConditional variational autoen- coder with adversarial learning for end-to-end text-to-speech,\u201d in ICML, 2021. [158] E. Casanova, J. Weber, C. D. Shulby, et al., \u201cYourTTS: Towards zero-shot multi-speaker TTS and zero-shot voice conversion for", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2209 }, { "text": "everyone,\u201d in ICML, 2022. [159] X. Tan, J. Chen, H. Liu, et al., \u201cNaturalSpeech: End-to-End Text-to- Speech Synthesis With Human-Level Quality,\u201d IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024. [160] S.-g. Lee, W. Ping, B. Ginsburg, et al., \u201cBigVGAN: A Universal Neural Vocoder with Large-Scale Training,\u201d in ICLR, 2023.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2210 }, { "text": "are zero-shot text to speech synthesizers,\u201d arXiv:2301.02111, 2023. [164] E. Kharitonov, D. Vincent, Z. Borsos, et al., \u201cSpeak, read and prompt: High-fidelity text-to-speech with minimal supervision,\u201d Transactions of the Association for Computational Linguistics, 2023. [165] D. Yang, J. Tian, X. Tan, et al., \u201cUniAudio: An Audio Foundation", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2211 }, { "text": "Model Toward Universal Audio Generation,\u201d arXiv:2310.00704, 2023. [166] R. Huang, M. W. Lam, J. Wang, et al., \u201cFastDiff: A fast conditional diffusion model for high-quality speech synthesis,\u201d in IJCAI, 2022. [167] Y. Ren, C. Hu, X. Tan, et al., \u201cFastSpeech 2: Fast and High-Quality End-to-End Text to Speech,\u201d in ICLR, 2020.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2212 }, { "text": "Speech Synthesis with Factorized Codec and Diffusion Models,\u201d arXiv:2403.03100, 2024. [171] C. Du, Y. Guo, F. Shen, et al., \u201cUniCATS: A unified context- aware text-to-speech framework with contextual vq-diffusion and vocoding,\u201d in AAAI, 2024. [172] D. Yang, S. Liu, R. Huang, et al., \u201cInstructTTS: Modelling Expres-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2213 }, { "text": "sive TTS in Discrete Latent Space with Natural Language Style Prompt,\u201d IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2024. [173] Y. Liu, M. Ott, N. Goyal, et al., \u201cRoBERTa: A robustly optimized BERT pretraining approach,\u201d arXiv:1907.11692, 2019. [174] Y. Jang, J.-H. Kim, J. Ahn, et al., \u201cFaces that speak: Jointly", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2214 }, { "text": "synthesising talking face and speech from text,\u201d in CVPR, 2024. [175] K. Cheng, X. Cun, Y. Zhang, et al., \u201cVideoReTalking: Audio-based Lip Synchronization for Talking Head Video Editing In the Wild,\u201d in SIGGRAPH, 2022. [176] S. Wang, L. Li, Y. Ding, et al., \u201cOne-shot talking face generation from single-speaker audio-visual correlation learning,\u201d in AAAI,", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2215 }, { "text": "IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. 50, NO. 1, NOVEMBER 2024 18 [177] J. Ling, X. Tan, L. Chen, et al., \u201cStableFace: Analyzing and Improving Motion Stability for Talking Face Generation,\u201d IEEE Journal of Selected Topics in Signal Processing, 2023. [178] Y. Gan, Z. Yang, X. Yue, et al., \u201cEfficient emotional adaptation for", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2216 }, { "text": "audio-driven talking-head generation,\u201d in ICCV, 2023. [179] H. Wei, Z. Yang, and Z. Wang, \u201cAniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation,\u201d arXiv:2403.17694, 2024. [180] C. Wang, K. Tian, J. Zhang, et al., \u201cV-Express: Conditional Dropout for Progressive Training of Portrait Video Generation,\u201d", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2217 }, { "text": "arXiv:2406.02511, 2024. [181] Z. Chen, J. Cao, Z. Chen, et al., \u201cEchoMimic: Lifelike Audio- Driven Portrait Animations through Editable Landmark Condi- tions,\u201d arXiv:2407.08136, 2024. [182] S. Xu, G. Chen, Y.-X. Guo, et al., \u201cVASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time,\u201d arXiv:2404.10667, 2024.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2218 }, { "text": "video generation,\u201d in NeurIPS, 2022. [190] Y. Lu, J. Chai, and X. Cao, \u201cLive Speech Portraits: Real-Time Photorealistic Talking-Head Animation,\u201d ACM Transactions on Graphics, 2021. [191] S. Gururani, A. Mallya, T.-C. Wang, et al., \u201cSpace: Speech-driven portrait animation with controllable expression,\u201d in ICCV, 2023.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2219 }, { "text": "Head Pose and Facial Expressions,\u201d in ICASSP, 2023. [194] Z. Peng, H. Wu, Z. Song, et al., \u201cEmoTalk: Speech-Driven Emo- tional Disentanglement for 3D Face Animation,\u201d in ICCV, 2023. [195] W. Zhong, C. Fang, Y. Cai, et al., \u201cIdentity-preserving talking face generation with landmark and appearance priors,\u201d in CVPR,", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2220 }, { "text": "2023. [196] Z. Yan, Y. Luo, S. Lyu, et al., \u201cTranscending Forgery Specificity with Latent Space Augmentation for Generalizable Deepfake Detection,\u201d in CVPR, 2024. [197] C. Tan, H. Liu, Y. Zhao, et al., \u201cRethinking the Up-Sampling Operations in CNN-Based Generative Network for Generalizable Deepfake Detection,\u201d in CVPR, 2024.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2221 }, { "text": "tion: Improving generalizability through frequency space domain learning,\u201d in AAAI, 2024. [203] B. M. Le and S. S. Woo, \u201cADD: Frequency attention and multi- view based knowledge distillation to detect low-quality com- pressed deepfake images,\u201d in AAAI, 2022. [204] M. Kim, S. Tariq, and S. S. Woo, \u201cFReTAL: Generalizing Deep-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2222 }, { "text": "fake Detection using Knowledge Distillation and Representation Learning,\u201d in CVPRW, 2021. [205] Y. Xu, K. Raja, L. Verdoliva, et al., \u201cLearning pairwise interaction for generalizable deepfake detection,\u201d in WACVW, 2023. [206] M. Du, S. Pentyala, Y. Li, et al., \u201cTowards generalizable deepfake detection with locality-aware autoencoder,\u201d in CIKM, 2019.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2223 }, { "text": "open-world deepfake attribution,\u201d in ICCV, 2023. [213] T. Zhao, X. Xu, M. Xu, et al., \u201cLearning self-consistency for deepfake detection,\u201d in ICCV, 2021. [214] A. Hooda, N. Mangaokar, R. Feng, et al., \u201cD4: Detection of adversarial diffusion deepfakes using disjoint ensembles,\u201d in WACV, 2024. [215] Y. Ju, S. Hu, S. Jia, et al., \u201cImproving fairness in deepfake detec-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2224 }, { "text": "tion,\u201d in WACV, 2024. [216] D. Tantaru, E. Oneata, and D. Oneata, \u201cWeakly-supervised deep- fake localization in diffusion-generated images,\u201d in WACV, 2024. [217] L. Trinh, M. Tsang, S. Rambhatla, et al., \u201cInterpretable and trust- worthy deepfake detection via dynamic prototypes,\u201d in WACV, 2021. [218] Z. Ba, Q. Liu, Z. Liu, et al., \u201cExposing the deception: Uncovering", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2225 }, { "text": "more forgery clues for deepfake detection,\u201d in AAAI, 2024. [219] T. Yang, Z. Huang, J. Cao, et al., \u201cDeepfake network architecture attribution,\u201d in AAAI, 2022. [220] Y. Nirkin, L. Wolf, Y. Keller, et al., \u201cDeepfake detection based on discrepancies between faces and their context,\u201d IEEE Transactions", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2226 }, { "text": "on Pattern Analysis and Machine Intelligence, 2022. [221] R. Lanzino, F. Fontana, A. Diko, et al., \u201cFaster than lies: Real-time deepfake detection using binary neural networks,\u201d in CVPRW, 2024. [222] A. Ciamarra, R. Caldelli, F. Becattini, et al., \u201cDeepfake detec- tion by exploiting surface anomalies: The surfake approach,\u201d in", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2227 }, { "text": "WACVW, 2024. [223] Y. Jeong, D. Kim, S. Min, et al., \u201cBiHPF: Bilateral High-Pass Filters for Robust Deepfake Detection,\u201d in WACV, 2022. [224] M. Tan, \u201cEfficientNet: Rethinking Model Scaling for Convolu- tional Neural Networks,\u201d in ICML, 2019. [225] F. Chollet, \u201cXception: Deep Learning With Depthwise Separable", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2228 }, { "text": "Convolutions,\u201d in CVPR, 2017. [226] K. He, X. Zhang, S. Ren, et al., \u201cDeep residual learning for image recognition,\u201d in CVPR, 2015. [227] Z. Liu, X. Qi, and P. H. Torr, \u201cGlobal texture enhancement for fake face detection in the wild,\u201d in CVPR, 2020. [228] Y. Luo, Y. Zhang, J. Yan, and W. Liu, \u201cGeneralizing face forgery", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2229 }, { "text": "detection with high-frequency features,\u201d in CVPR, 2021. [229] Y. Qian, G. Yin, L. Sheng, et al., \u201cThinking in frequency: Face forgery detection by mining frequency-aware clues,\u201d in ECCV, 2020. [230] S. Hussain, P. Neekhara, M. Jere, et al., \u201cAdversarial deepfakes: Evaluating vulnerability of deepfake detectors to adversarial", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2230 }, { "text": "examples,\u201d in WACV, 2021. [231] Z. Yang, J. Liang, Y. Xu, et al., \u201cMasked relation learning for deepfake detection,\u201d IEEE Transactions on Information Forensics and Security, 2023. [232] Y. Wang, K. Yu, C. Chen, et al., \u201cDynamic graph learning with content-guided spatial-frequency relation reasoning for deepfake", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2231 }, { "text": "detection,\u201d in CVPR, 2023. [233] J. Cao, C. Ma, T. Yao, et al., \u201cEnd-to-end reconstruction- classification learning for face forgery detection,\u201d in CVPR, 2022. [234] A. Dosovitskiy, L. Beyer, A. Kolesnikov, et al., \u201cAn image is worth 16x16 words: Transformers for image recognition at scale,\u201d in ICLR, 2021.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2232 }, { "text": "IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. 50, NO. 1, NOVEMBER 2024 19 [237] S. Kamat, S. Agarwal, T. Darrell, et al., \u201cRevisiting generalizability in deepfake detection: Improving metrics and stabilizing trans- fer,\u201d in ICCVW, 2023. [238] Y. Jeong, D. Kim, Y. Ro, et al., \u201cFrePGAN: Robust Deepfake", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2233 }, { "text": "Detection Using Frequency-Level Perturbations,\u201d in AAAI, 2022. [239] F. Lugstein, S. Baier, G. Bachinger, et al., \u201cPRNU-based Deepfake Detection,\u201d in IH&MMSec, 2021. [240] L. Guarnera, O. Giudice, and S. Battiato, \u201cDeepfake detection by analyzing convolutional traces,\u201d in CVPRW, 2020. [241] S. Agarwal, H. Farid, T. El-Gaaly, et al., \u201cDetecting Deep-Fake", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2234 }, { "text": "Videos from Appearance and Behavior,\u201d in WIFS, 2020. [242] D. Cozzolino, A. R\u00a8ossler, J. Thies, et al., \u201cID-Reveal: Identity- aware DeepFake Video Detection,\u201d in ICCV, 2021. [243] Z. Gu, Y. Chen, T. Yao, et al., \u201cDelving into the Local: Dynamic Inconsistency Learning for DeepFake Video Detection,\u201d in AAAI,", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2235 }, { "text": "2022. [244] Z. Gu, T. Yao, Y. Chen, et al., \u201cHierarchical Contrastive Inconsis- tency Learning for Deepfake Video Detection,\u201d in ECCV, 2022. [245] Y. Zhao, W. Ge, W. Li, et al., \u201cCapturing the Persistence of Facial Expression Features for Deepfake Video Detection,\u201d in ICICS, 2020. [246] A. Haliassos, K. Vougioukas, S. Petridis, et al., \u201cLips don\u2019t lie: A", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2236 }, { "text": "generalisable and robust approach to face forgery detection,\u201d in CVPR, 2021. [247] A. Haliassos, R. Mira, S. Petridis, et al., \u201cLeveraging real talking faces via self-supervision for robust forgery detection,\u201d in CVPR, 2022. [248] I. Demir and U. A. C\u00b8 iftc\u00b8i, \u201cHow Do Deepfakes Move? Motion Magnification for Deepfake Source Detection,\u201d in WACV, 2024.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2237 }, { "text": "Detection with Automatic Face Weighting,\u201d in CVPR, 2020. [252] I. Amerini and R. Caldelli, \u201cExploiting Prediction Error Incon- sistencies through LSTM-based Classifiers to Detect Deepfake Videos,\u201d in IH&MMSec, 2020. [253] I. Masi, A. Killekar, R. M. Mascarenhas, et al., \u201cTwo-branch Recurrent Network for Isolating Deepfakes in Videos,\u201d in ECCV,", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2238 }, { "text": "2020. [254] J. Deng, J. Guo, N. Xue, et al., \u201cArcFace: Additive Angular Margin Loss for Deep Face Recognition,\u201d in CVPR, 2019. [255] N. Bonettini, E. D. Cannas, S. Mandelli, et al., \u201cVideo face manip- ulation detection through ensemble of CNNs,\u201d in ICPR, 2021. [256] T. Wang and K. P. Chow, \u201cNoise Based Deepfake Detection via", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2239 }, { "text": "Multi-Head Relative-Interaction,\u201d in AAAI, 2023. [257] Z. Yan, Y. Zhao, S. Chen, et al., \u201cGeneralizing Deepfake Video Detection with Plug-and-Play: Video-Level Blending and Spa- tiotemporal Adapter Tuning,\u201d arXiv:2408.17065, 2024. [258] A. Coccomini, N. Messina, C. Gennaro, et al., \u201cCombining Effi-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2240 }, { "text": "cientNet and Vision Transformers for Video Deepfake Detection,\u201d in ICIAP, 2022. [259] A. Vaswani, N. Shazeer, N. Parmar, et al., \u201cAttention is all you need,\u201d in NeurIPS, 2017. [260] J. Guan, H. Zhou, Z. Hong, et al., \u201cDelving into sequential patches for deepfake detection,\u201d in NeurIPS, 2022. [261] E. Richardson, Y. Alaluf, O. Patashnik, et al., \u201cEncoding in Style:", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2241 }, { "text": "a StyleGAN Encoder for Image-to-Image Translation,\u201d in CVPR, 2021. [262] L. Tan, Y. Wang, J. Wang, et al., \u201cDeepfake Video Detection via Facial Action Dependencies Estimation,\u201d in AAAI, 2023. [263] Y. Xu, J. Liang, G. Jia, et al., \u201cTALL: Thumbnail Layout for Deepfake Video Detection,\u201d in ICCV, 2023.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2242 }, { "text": "TCN and orthogonal regularization for fake audio detection,\u201d in INTERSPEECH, 2023. [267] E. Conti, D. Salvi, C. Borrelli, et al., \u201cDeepfake speech detection through emotion recognition: a semantic approach,\u201d in ICASSP, 2022. [268] G. Hua, A. B. J. Teoh, and H. Zhang, \u201cTowards end-to-end synthetic speech detection,\u201d IEEE Signal Processing Letters, 2021.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2243 }, { "text": "using convolutional transformer-based spectrogram analysis,\u201d in ACSSC, 2021. [274] E. R. Bartusiak and E. J. Delp, \u201cTransformer-based speech syn- thesizer attribution in an open set scenario,\u201d in ICMLA, 2022. [275] J. M. Mart\u00b4\u0131n-Do\u02dcnas and A. \u00b4Alvarez, \u201cThe Vicomtech Audio Deepfake Detection System Based on Wav2vec2 for the 2022 ADD", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2244 }, { "text": "Challenge,\u201d in ICASSP, 2022. [276] Z. Cai, W. Wang, and M. Li, \u201cWaveform boundary detection for partially spoofed audio,\u201d in ICASSP, 2023. [277] Z. Zhang, X. Yi, and X. Zhao, \u201cFake speech detection using resid- ual network with transformer encoder,\u201d in IH&MMSec, 2021. [278] X. Liu, M. Liu, L. Wang, et al., \u201cLeveraging positional-related", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2245 }, { "text": "local-global dependency for synthetic speech detection,\u201d in ICASSP, 2023. [279] R. Wang, F. Juefei-Xu, Y. Huang, et al., \u201cDeepSonar: Towards Effective and Robust Detection of AI-Synthesized Fake Voices,\u201d in ACMMM, 2020. [280] X. Zhang, J. Yi, C. Wang, et al., \u201cWhat to remember: Self-adaptive continual learning for audio deepfake detection,\u201d in AAAI, 2024.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2246 }, { "text": "to-end Dense Swin Transformer deep learning model for audio\u2013 visual deepfakes detection,\u201d Applied Soft Computing, 2023. [286] S. Asha, P. Vinod, and V. G. Menon, \u201cA defensive attention mech- anism to detect deepfake content across multiple modalities,\u201d Multimedia Systems, 2024. [287] X. Liu, Y. Yu, X. Li, et al., \u201cMagnifying multimodal forgery clues", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2247 }, { "text": "for deepfake detection,\u201d Signal Processing: Image Communication, 2023. [288] C. Feng, Z. Chen, and A. Owens, \u201cSelf-supervised video forensics by audio-visual anomaly detection,\u201d in CVPR, 2023. [289] H. Zou, M. Shen, Y. Hu, et al., \u201cCross-modality and within- modality regularization for audio-visual deepfake detection,\u201d in", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2248 }, { "text": "ICASSP, 2024. [290] F. Nie, J. Ni, J. Zhang, et al., \u201cFRADE: Forgery-aware Audio- distilled Multimodal Learning for Deepfake Detection,\u201d in ACMMM, 2024. [291] Y. Zhang, W. Lin, and J. Xu, \u201cJoint audio-visual attention with contrastive learning for more general deepfake detection,\u201d ACM Transactions on Multimedia Computing, Communications and Appli-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2249 }, { "text": "cations, 2024. [292] H. Dang, F. Liu, J. Stehouwer, et al., \u201cOn the detection of digital face manipulation,\u201d in CVPR, 2020. [293] R. Wang, F. Juefei-Xu, L. Ma, et al., \u201cFakeSpotter: A Simple yet Robust Baseline for Spotting AI-Synthesized Fake Faces,\u201d in IJCAI, 2020. [294] Y. He, B. Gan, S. Chen, et al., \u201cForgeryNet: A Versatile Benchmark", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2250 }, { "text": "for Comprehensive Forgery Analysis,\u201d in CVPR, 2021. [295] Z. Chen, K. Sun, Z. Zhou, et al., \u201cDiffusionFace: Towards a Com- prehensive Dataset for Diffusion-Based Face Forgery Analysis,\u201d arXiv:2403.18471, 2024. [296] L. Verdoliva, C. Riess, J. Thies, et al., \u201cFaceForensics++: Learning to Detect Manipulated Facial Images,\u201d in ICCV, 2019.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2251 }, { "text": "IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. 50, NO. 1, NOVEMBER 2024 20 [298] Y. Li, X. Yang, P. Sun, et al., \u201cCeleb-DF: A Large-scale Challenging Dataset for DeepFake Forensics,\u201d in CVPR, 2020. [299] B. Zi, M. Chang, J. Chen, et al., \u201cWildDeepfake: A Challenging Real-World Dataset for Deepfake Detection,\u201d in ACMMM, 2020.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2252 }, { "text": "mark,\u201d arXiv:2405.19707, 2024. [303] J. Frank and L. Sch\u00a8onherr, \u201cWaveFake: A Data Set to Facilitate Audio Deepfake Detection,\u201d in NeurIPS, 2021. [304] X. Wang, J. Yamagishi, M. Todisco, et al., \u201cASVspoof 2019: A large-scale public database of synthesized, converted and re- played speech,\u201d Computer Speech & Language, 2020.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2253 }, { "text": "Synthesis Detection Challenge,\u201d in ICASSP, 2022. [308] J. Yi, J. Tao, R. Fu, et al., \u201cADD 2023: the Second Audio Deepfake Detection Challenge,\u201d arXiv:2305.13774, 2023. [309] R. Reimao and V. Tzerpos, \u201cFoR: A Dataset for Synthetic Speech Detection,\u201d in SpeD, 2019. [310] N. M\u00a8uller, P. Kawa, W. Choong, et al., \u201cMLAAD: The Multi-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2254 }, { "text": "Language Audio Anti-Spoofing Dataset,\u201d IJCNN, 2024. [311] H. Khalid, S. Tariq, M. Kim, et al., \u201cFakeAVCeleb: A novel audio- video multimodal deepfake dataset,\u201d in NeurIPS, 2021. [312] Z. Cai, K. Stefanov, A. Dhall, et al., \u201cDo You Really Mean That? Content Driven Audio-Visual Deepfake Dataset and Multimodal", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2255 }, { "text": "Method for Temporal Forgery Localization,\u201d in DICTA, 2022. [313] B. Dolhansky, J. Bitton, B. Pflaum, et al., \u201cThe DeepFake Detection Challenge (DFDC) Dataset,\u201d arXiv:2006.07397, 2020. [314] W. Yang, X. Zhou, Z. Chen, et al., \u201cAVoiD-DF: Audio-Visual Joint Learning for Detecting Deepfake,\u201d IEEE Transactions on", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2256 }, { "text": "Information Forensics and Security, 2023. [315] B. Hosler, D. Salvi, A. Murray, et al., \u201cDo deepfakes feel emo- tions? a semantic approach to detecting deepfakes via emotional inconsistencies,\u201d in CVPR, 2021. [316] H. Tak, M. Todisco, X. Wang, et al., \u201cAutomatic speaker verifica- tion spoofing and deepfake detection using wav2vec 2.0 and data", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2257 }, { "text": "augmentation,\u201d arXiv:2202.12233, 2022. [317] V. S. Katamneni and A. Rattani, \u201cMIS-AVoiDD: Modality in- variant and specific representation for audio-visual deepfake detection,\u201d in ICMLA, 2023. [318] Y. Yu, X. Liu, R. Ni, et al., \u201cPVASS-MDD: Predictive visual-audio alignment self-supervision for multimodal deepfake detection,\u201d", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2258 }, { "text": "IEEE Transactions on Circuits and Systems for Video Technology, 2023. [319] T. Liu, F. Chen, F. S., et al., \u201cAniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion Encoding,\u201d arXiv:2405.03121, 2024. [320] K. Cho, J. Lee, H. Yoon, et al., \u201cGaussianTalker: Real-Time Talking", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2259 }, { "text": "Head Synthesis with 3D Gaussian Splatting,\u201d in ACMMM, 2024. [321] Y. Zheng, H. Yang, T. Zhang, et al., \u201cGeneral facial representation learning in a visual-linguistic manner,\u201d in CVPR, 2022. [322] Z. Zhang, L. Li, Y. Ding, et al., \u201cFlow-guided one-shot talking face generation with a high-resolution audio-visual dataset,\u201d in", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2260 }, { "text": "CVPR, 2021. [323] I. Demirsahin, O. Kjartansson, A. Gutkin, et al., \u201cOpen-source multi-speaker corpora of the English accents in the British isles,\u201d in LREC, 2020. [324] Y. A. Li, C. Han, V. Raghavan, et al., \u201cStyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adver- sarial Training with Large Speech Language Models,\u201d in NeurIPS,", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2261 }, { "text": "2023. [325] H. Wang, M. Yu, J. Hai, et al., \u201cSSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis,\u201d arXiv:2409.07556, 2024. [326] H. Zen, V. Dang, R. Clark, et al., \u201cLibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech,\u201d in INTERSPEECH, 2019. [327] T.-C. Wang, A. Mallya, and M.-Y. Liu, \u201cOne-shot free-view neural", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2262 }, { "text": "talking-head synthesis for video conferencing,\u201d in CVPR, 2021. [328] Z. Yan, Y. Zhang, X. Yuan, et al., \u201cDeepfakeBench: A Comprehen- sive Benchmark of Deepfake Detection,\u201d in NeurIPS, 2023. [329] D. P. Kingma and M. Welling, \u201cAuto-Encoding Variational Bayes,\u201d arXiv, 2013. [330] J. Ho, A. Jain, and P. Abbeel, \u201cDenoising diffusion probabilistic", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2263 }, { "text": "models,\u201d in NeurIPS, 2020. Florinel-Alin Croitoru is a PhD student at the University of Bucharest, Romania. In 2021, he obtained his masters degree in Artificial Intelligence with a thesis on action spotting in football videos. His domains of interest include machine learning, computer vision and deep learning. He", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2264 }, { "text": "published several studies in top-tier conferences and jour- nals, such as CVPR, ECAI, TPAMI, IJCV, CVIU. Andrei H\u02c6\u0131ji is a PhD student at the University of Bucharest, Romania. He obtained his master\u2019s degree in Information Security from the Military Technical Academy \u201cFerdinand I\u201d in 2019. His research interests include machine learning,", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2265 }, { "text": "anomaly detection and cybersecurity. Vlad Hondru is a PhD student at the University of Bucharest, Romania. He obtained his bachelor\u2019s degree from the University of Manchester in Mechatronic Engi- neering, then he graduated from Imperial College London, studying towards an MSc in Computing Science, with a", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2266 }, { "text": "Visual Computing and Robotics specialization, focusing on Artificial Intelligence. Nicolae-C\u02d8at\u02d8alin Ristea graduated as valedictorian from the Faculty of Electronics, Telecommunications and Infor- mation Technology, NUST Politehnica Bucharest, in 2019. He completed his PhD in 2024 at the same university.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2267 }, { "text": "Nicolae is co-author of multiple papers accepted at top- tier conferences and journals, such as CVPR, WACV and TPAMI. His research interests include deep learning, com- puter vision, machine learning and signal processing. Paul Irofti is an Associate Professor within the Computer Science Department of the Faculty of Mathematics and", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2268 }, { "text": "Computer Science at the University of Bucharest. He is the co-author of the book \u201cDictionary Learning Algorithms and Applications\u201d (Springer 2018). He is PhD in Systems Engineering at the Politehnica University of Bucharest since 2016. His interests are anomaly detection, signal processing, numerical algorithms and optimization.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2269 }, { "text": "Marius Popescu is associate professor at the University of Bucharest, Romania. He defended his PhD in 2004. His domains of interest are: AI, ML, computational linguistics, computer vision. His achievements in these fields include methods that ranked 3rd in the NLI Shared Task of BEA-8, 4th in the FER Challenge of WREPL 2013, 2nd in the ADI", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2270 }, { "text": "Shared Task of VarDial 2016, 1st in the NLI Shared Task of BEA-12. Cristian Rusu is associate professor within the Computer Science Department of the Faculty of Mathematics and Computer Science at the University of Bucharest. He re- ceived his PhD in Systems Engineering at the Politehnica University of Bucharest in 2012. His research interests in-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2271 }, { "text": "clude signal processing, numerical linear algebra, machine learning, and deep learning. Radu Ionescu is full professor at the University of Bucharest, Romania. He completed his PhD at the Univer- sity of Bucharest in 2013. His research interests include machine learning, computer vision, image processing,", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2272 }, { "text": "IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. 50, NO. 1, NOVEMBER 2024 21 Fahad Khan is a faculty member at MBZ University of AI (MBZUAI), UAE and Link\u00a8oping University, Sweden. He received the M.Sc. degree in Intelligent Systems Design from Chalmers University of Technology, Sweden and a", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2273 }, { "text": "Ph.D. degree in Computer Vision from Autonomous Uni- versity of Barcelona, Spain. His research interests include a wide range of topics within computer vision, such as ob- ject recognition, object detection, action recognition and visual tracking. Mubarak Shah is the UCF Trustee chair professor and the", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2274 }, { "text": "founding director of the Center for Research in Computer Vision at the University of Central Florida (UCF). He is a fellow of the NAI, IEEE, AAAS, IAPR and SPIE. He is an editor of an international book series on video computing, was editor-in-chief of Machine Vision and Applications and an associate editor of ACM Computing Surveys and IEEE", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2275 }, { "text": "IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. 50, NO. 1, NOVEMBER 2024 22 A I xsxs A(xt) xt I(xs) G \u03b8 D \u03a6 I x\u2019 LG /LD I(x\u2019) Identity-specific loss Fig. 4. An overview of face swapping based on GANs. The generative process is conditioned on an identity encoder I and an attribute encoder", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2276 }, { "text": "A, aiming to preserve the target attributes while replacing the target identity with the source identity. 8 SUPPLEMENTARY: DEEPFAKE GENERATION TU- TORIAL Various deep generative models are actively being used to successfully generate deepfake content. Among the deep- fake generative methods, we next explain in detail how the", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2277 }, { "text": "most popular and interesting developments work, such as generative adversarial networks, variational autoencoders, diffusion models, as well as Neural Radiance Fields. 8.1 Generative Adversarial Networks Generative adversarial networks (GANs) [64] consist of two neural networks, called the generator and the discriminator.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2278 }, { "text": "The generator, denoted as G\u03b8(z), transforms input Gaussian noise z \u223cp(z) into a sample from the data distribution. The discriminator, represented as D\u03d5(x), outputs a single scalar value that predicts the probability of a given sample x to be real, rather than being generated by G\u03b8. Therefore, the discriminator D\u03d5 is trained as a binary classifier, where the", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2279 }, { "text": "real samples are labeled with 1 and the generated ones with 0, as follows: min D\u03d5 LD = \u2212Ex\u223cp(x)[log (D\u03d5(x))] \u2212Ez\u223cp(z)[log (1\u2212D\u03d5(G\u03b8(z)))] , (1) where p(x) and p(z) represent the real data distribution and the Gaussian distribution, respectively, with p(z) serving as a source for sampling varied inputs for the generator. The", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2280 }, { "text": "generator is trained to deceive the discriminator. Thus, its training objective maximizes the probability assigned by the discriminator to the generated samples D\u03d5(G\u03b8(z)): min G\u03b8 LG = \u2212Ez\u223cp(z) [log (D\u03d5(G\u03b8(z)))]. (2) We can also express the whole training framework through a single mini-max optimization objective, as follows:", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2281 }, { "text": "min G\u03b8 max D\u03d5 Ex\u223cp(x)[log(D\u03d5(x))]+Ez\u223cp(z)[log(1\u2212D\u03d5(G\u03b8(z)))], (3) Note that optimizing for G\u03b8 does not influence the first term of the objective, as it depends only on D\u03d5. Usage example. In Figure 4, we illustrate a typical face swapping pipeline powered by GANs. The generator G\u03b8 is conditioned on features derived from two sources. The", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2282 }, { "text": "identity encoder I extracts features from the identity source image xs, while the attribute encoder A extracts features from the target image xt. Commonly, the encoder I is E\u03a6 \u03bcx \u03c3x \u03b5 ~ p(z) x z = \u03bcx + \u03c3x\u03b5 D\u03b8 ||D\u03b8(z)-x||2 D\u03b8(z) KL(E\u03a6(z|x)||p(z)) Fig. 5. An overview of face synthesis based on VAEs. The KL divergence", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2283 }, { "text": "is used to minimize the distribution gap between the distribution of z and the standard Gaussian distribution p(z). a (pre-trained) face recognition model, while A is a ran- domly initialized encoder trained along with the rest of the pipeline. The generator harnesses the input features to produce an image x\u2032 that retains the attributes of xt, while", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2284 }, { "text": "swapping the target identity with the source identity from xs. Aside from the previously discussed adversarial losses, the pipeline incorporates a reconstruction loss between xt and x\u2032 to ensure attribute preservation. Additionally, an identity-specific loss is employed, typically calculated as the", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2285 }, { "text": "cosine similarity between the feature vectors extracted by the identity encoder I from x\u2032 and xs. 8.2 Variational Autoencoders A Variational Autoencoder (VAE) [329] is a modified version of the classic autoencoder, designed to support generative modeling by learning a probabilistic latent space. In the", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2286 }, { "text": "classic autoencoder case, during training, an encoder maps a sample x to a latent representation z from which a decoder is tasked to reconstruct the original input x. This approach is unsuitable for generative modeling because z follows an arbitrary complex distribution, making it impossible to directly sample from it. Kingma et al. [329] address this", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2287 }, { "text": "limitation by enforcing z to follow a standard Gaussian dis- tribution. They achieved this by adding a Kullback-Leibler (KL) divergence regularization term to the loss function, which minimizes the divergence between the distribution of z and the standard Gaussian distribution. To model the dis- tribution of z, the encoder outputs a mean \u00b5x and a variance", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2288 }, { "text": "\u03c3x that describe a Gaussian distribution. During training z is sampled from this distribution. The loss function is as follows: L = Ez\u223cE\u03d5(z|x) \u0002\u2225D\u03b8(z) \u2212x\u22252 2 \u0003 + KL (E\u03d5(z|x)\u2225p(z)) , (4) where E\u03d5(z|x) denotes the encoder, D\u03b8(x|z) is the decoder and p(z) represents the standard Gaussian distribution.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2289 }, { "text": "Usage example. In Figure 5, we showcase the training process of a VAE for face synthesis. To enable gradient back- propagation through the encoder, the reparameterization trick is applied to sample z \u223cN(\u00b5x, \u03c3xI). During inference, z is directly sampled from p(z) and fed into the decoder. 8.3 Diffusion Models", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2290 }, { "text": "Diffusion models [1] consist of two diffusion processes, the forward diffusion process and the reverse diffusion process. In the forward process, a data sample is progressively transformed over T steps by adding Gaussian noise at each step, eventually converting it into an approximate standard Gaussian distribution. The reverse process operates in the", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2291 }, { "text": "IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. 50, NO. 1, NOVEMBER 2024 23 Image Encoder An image of a woman smiling Text Encoder x0 \u03b5 xt \u03b5\u03b8(xt, t) \u03b5 Lsimple Forward process Training Generation Image Encoder An image of a man dressed as superman Text Encoder \u03b5 Reverse process Fig. 6. An overview of text-conditional personalized generation based", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2292 }, { "text": "on diffusion models. The model aims to generate images of a source identity conditioned on a text prompt. opposite direction, serving as the generative mechanism. It learns to map a standard Gaussian sample back to a data sample. Forward Process. The forward process is a Markov chain xt \u223cq(xt|xt\u22121) = N(\u221a1 \u2212\u03b2t\u00b7xt\u22121, \u03b2t\u00b7I), which gradually", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2293 }, { "text": "adds Gaussian noise that depends on a variance schedule {\u03b2t}T t=1, where x0 \u223cq(x0) represents a data sample and xT is approximately a standard Gaussian sample. This for- mulation supports an efficient sampling for an arbitrary xt during training: xt \u223cN(\u221a\u00af\u03b1tx0, (1 \u2212\u00af\u03b1t) \u00b7 I), (5) where \u03b1t = 1 \u2212\u03b2t, \u00af\u03b1t = Qt", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2294 }, { "text": "i=0 \u03b1i. Reverse Process. The reverse process starts with xT \u223c N(0, I) and follows the learned Gaussian transitions de- noted as p\u03b8(xt\u22121|xt) = N(\u00b5\u03b8(xt, t), \u03a3\u03b8(xt, t)) to recover a data sample x0. Commonly, in practice, the variance \u03a3\u03b8(xt, t) is approximated with \u03b2t. Thus, the only learnable component that remains is the mean \u00b5\u03b8(xt, t), which can be", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2295 }, { "text": "rewritten as a function of noise [330]: \u00b5\u03b8(xt, t) = 1 \u221a\u03b1t \u0012 xt \u2212 \u03b2t \u221a1 \u2212\u00af\u03b1t \u00b7 \u03f5\u03b8(xt, t) \u0013 . (6) As a consequence, during training, the neural network, denoted by \u03f5\u03b8(xt, t), learns to approximate the noise \u03f5 \u223c N(0, I) added at arbitrary steps t \u223cU(1, . . . , T) to the data samples x0 \u223cq(x0). Formally, the optimization objective for", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2296 }, { "text": "the reverse process is defined as: min \u03f5\u03b8 Lsimple = Ex0\u223cq(x0),t\u223cU(1,...,T ),\u03f5\u223cN(0,I)\u2225\u03f5 \u2212\u03f5\u03b8(xt, t)\u22252 2. (7) Usage example. In Figure 6, we illustrate the training and inference processes of a text-conditional personalized gener- ative pipeline for a given identity. During training, pairs of images representing the same identity are used. One image", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2297 }, { "text": "undergoes the forward process, and the model is tasked with predicting the added noise. The second image in the Pose: R, t Audio Encoder c \u03c3 Rendering Talking head video Audio signal Training frames Fig. 7. An overview of talking-head synthesis based on NeRFs. The model learns to predict color and density values, which furtherenable", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2298 }, { "text": "the rendering of the deepfake talking head video. pair, along with a text description of the first image, serve as conditional inputs to guide the model. The generative (reverse) process begins with standard Gaussian noise and progressively generates an image that aligns with the pro- vided text description, while preserving the identity from", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2299 }, { "text": "the given source image. 8.4 Neural Radiance Fields Neural Radiance Fields (NeRFs) [151] represent a method for synthesizing novel views from a sparse set of input images of a scene. The core idea is to train a single neural network to overfit to a specific scene, with the weights en- coding the detailed information and structure of the scene.", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2300 }, { "text": "The input to the neural network is a 5D vector consisting of three spatial coordinates (x, y, z), which represent a point in 3D space of the scene, and two angles (\u03b8, \u03d5), which define the viewing direction. The network outputs the density \u03c3 at the given spatial point and its color c, represented as an RGB vector. However, datasets typically lack direct", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2301 }, { "text": "annotations for the density and color of each spatial point, so the network optimization is performed in the pixel space instead. This process involves selecting a viewing direction and casting a ray through the scene, along which multiple spatial points are sampled. These points are processed by the neural network to compute their corresponding density", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2302 }, { "text": "and color values. The outputs are then combined using volume rendering to generate an approximate pixel value. This predicted pixel value is compared with the ground- truth pixel using a mean squared error loss function to optimize the network. Formally, the loss function is defined as: L = \u03a3r\u223cR\u2225\u02c6C(r) \u2212C(r)\u22252", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2303 }, { "text": "2, (8) where R is the set of rays in the current batch, \u02c6C(r) is the predicted color for a particular ray r and C(r) is the ground- truth color. In practice, NeRFs have two implementation details for a better optimization. First, spatial positions are projected into a higher-dimensional space using a positional encod-", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2304 }, { "text": "ing, similar to the approach used for vision transformers. Second, NeRFs use Hierarchical Volume Sampling, which involves training two neural networks. The first network samples coarse spatial points along a ray to estimate the overall structure of the scene. It then identifies regions of greater importance. The second network focuses on these", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2305 }, { "text": "regions, performing finer sampling to achieve more accurate results. Usage example. NeRFs are used for talking head synthesis, as illustrated in Figure 7. The main idea is to use the audio as the driving signal. More precisely, along with the infor- mation about the pose, the neural network processes audio", "source": "Deepfake Detection Generative AI Era", "year": 2024, "url": "https://arxiv.org/abs/2411.19537", "id": 2306 }, { "text": "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks Patrick Lewis\u2020\u2021, Ethan Perez\u22c6, Aleksandra Piktus\u2020, Fabio Petroni\u2020, Vladimir Karpukhin\u2020, Naman Goyal\u2020, Heinrich K\u00fcttler\u2020, Mike Lewis\u2020, Wen-tau Yih\u2020, Tim Rockt\u00e4schel\u2020\u2021, Sebastian Riedel\u2020\u2021, Douwe Kiela\u2020 \u2020Facebook AI Research; \u2021University College London; \u22c6New York University;", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2307 }, { "text": "plewis@fb.com Abstract Large pre-trained language models have been shown to store factual knowledge in their parameters, and achieve state-of-the-art results when \ufb01ne-tuned on down- stream NLP tasks. However, their ability to access and precisely manipulate knowl- edge is still limited, and hence on knowledge-intensive tasks, their performance", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2308 }, { "text": "lags behind task-speci\ufb01c architectures. Additionally, providing provenance for their decisions and updating their world knowledge remain open research problems. Pre- trained models with a differentiable access mechanism to explicit non-parametric memory have so far been only investigated for extractive downstream tasks. We", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2309 }, { "text": "explore a general-purpose \ufb01ne-tuning recipe for retrieval-augmented generation (RAG) \u2014 models which combine pre-trained parametric and non-parametric mem- ory for language generation. We introduce RAG models where the parametric memory is a pre-trained seq2seq model and the non-parametric memory is a dense", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2310 }, { "text": "vector index of Wikipedia, accessed with a pre-trained neural retriever. We com- pare two RAG formulations, one which conditions on the same retrieved passages across the whole generated sequence, and another which can use different passages per token. We \ufb01ne-tune and evaluate our models on a wide range of knowledge-", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2311 }, { "text": "intensive NLP tasks and set the state of the art on three open domain QA tasks, outperforming parametric seq2seq models and task-speci\ufb01c retrieve-and-extract architectures. For language generation tasks, we \ufb01nd that RAG models generate more speci\ufb01c, diverse and factual language than a state-of-the-art parametric-only", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2312 }, { "text": "seq2seq baseline. 1 Introduction Pre-trained neural language models have been shown to learn a substantial amount of in-depth knowl- edge from data [47]. They can do so without any access to an external memory, as a parameterized implicit knowledge base [51, 52]. While this development is exciting, such models do have down-", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2313 }, { "text": "sides: They cannot easily expand or revise their memory, can\u2019t straightforwardly provide insight into their predictions, and may produce \u201challucinations\u201d [38]. Hybrid models that combine parametric memory with non-parametric (i.e., retrieval-based) memories [20, 26, 48] can address some of these issues because knowledge can be directly revised and expanded, and accessed knowledge can be", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2314 }, { "text": "The Divine Comedy (x) q Query Encoder q(x) MIPS p\u03b8 Generator p\u03b8 (Parametric) Margin- alize This 14th century work is divided into 3 sections: \"Inferno\", \"Purgatorio\" & \"Paradiso\" (y) End-to-End Backprop through q and p\u03b8 Barack Obama was born in Hawaii.(x) Fact Veri\ufb01cation: Fact Query supports (y)", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2315 }, { "text": "Question Generation Fact Veri\ufb01cation: Label Generation Document Index Define \"middle ear\"(x) Question Answering: Question Query The middle ear includes the tympanic cavity and the three ossicles. (y) Question Answering: Answer Generation Retriever p\u03b7 (Non-Parametric) z4 z3 z2 z1 d(z) Jeopardy Question", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2316 }, { "text": "Generation: Answer Query Figure 1: Overview of our approach. We combine a pre-trained retriever (Query Encoder + Document Index) with a pre-trained seq2seq model (Generator) and \ufb01ne-tune end-to-end. For query x, we use Maximum Inner Product Search (MIPS) to \ufb01nd the top-K documents zi. For \ufb01nal prediction y, we", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2317 }, { "text": "treat z as a latent variable and marginalize over seq2seq predictions given different documents. but have only explored open-domain extractive question answering. Here, we bring hybrid parametric and non-parametric memory to the \u201cworkhorse of NLP,\u201d i.e. sequence-to-sequence (seq2seq) models. We endow pre-trained, parametric-memory generation models with a non-parametric memory through", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2318 }, { "text": "a general-purpose \ufb01ne-tuning approach which we refer to as retrieval-augmented generation (RAG). We build RAG models where the parametric memory is a pre-trained seq2seq transformer, and the non-parametric memory is a dense vector index of Wikipedia, accessed with a pre-trained neural retriever. We combine these components in a probabilistic model trained end-to-end (Fig. 1). The", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2319 }, { "text": "retriever (Dense Passage Retriever [26], henceforth DPR) provides latent documents conditioned on the input, and the seq2seq model (BART [32]) then conditions on these latent documents together with the input to generate the output. We marginalize the latent documents with a top-K approximation, either on a per-output basis (assuming the same document is responsible for all tokens) or a per-token", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2320 }, { "text": "basis (where different documents are responsible for different tokens). Like T5 [51] or BART, RAG can be \ufb01ne-tuned on any seq2seq task, whereby both the generator and retriever are jointly learned. There has been extensive previous work proposing architectures to enrich systems with non-parametric memory which are trained from scratch for speci\ufb01c tasks, e.g. memory networks [64, 55], stack-", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2321 }, { "text": "augmented networks [25] and memory layers [30]. In contrast, we explore a setting where both parametric and non-parametric memory components are pre-trained and pre-loaded with extensive knowledge. Crucially, by using pre-trained access mechanisms, the ability to access knowledge is present without additional training.", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2322 }, { "text": "Our results highlight the bene\ufb01ts of combining parametric and non-parametric memory with genera- tion for knowledge-intensive tasks\u2014tasks that humans could not reasonably be expected to perform without access to an external knowledge source. Our RAG models achieve state-of-the-art results on open Natural Questions [29], WebQuestions [3] and CuratedTrec [2] and strongly outperform", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2323 }, { "text": "recent approaches that use specialised pre-training objectives on TriviaQA [24]. Despite these being extractive tasks, we \ufb01nd that unconstrained generation outperforms previous extractive approaches. For knowledge-intensive generation, we experiment with MS-MARCO [1] and Jeopardy question generation, and we \ufb01nd that our models generate responses that are more factual, speci\ufb01c, and", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2324 }, { "text": "diverse than a BART baseline. For FEVER [56] fact veri\ufb01cation, we achieve results within 4.3% of state-of-the-art pipeline models which use strong retrieval supervision. Finally, we demonstrate that the non-parametric memory can be replaced to update the models\u2019 knowledge as the world changes.1 2 Methods", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2325 }, { "text": "We explore RAG models, which use the input sequence x to retrieve text documents z and use them as additional context when generating the target sequence y. As shown in Figure 1, our models leverage two components: (i) a retriever p\u03b7(z|x) with parameters \u03b7 that returns (top-K truncated) distributions over text passages given a query x and (ii) a generator p\u03b8(yi|x, z, y1:i\u22121) parametrized", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2326 }, { "text": "by \u03b8 that generates a current token based on a context of the previous i \u22121 tokens y1:i\u22121, the original input x and a retrieved passage z. To train the retriever and generator end-to-end, we treat the retrieved document as a latent variable. We propose two models that marginalize over the latent documents in different ways to produce a", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2327 }, { "text": "distribution over generated text. In one approach, RAG-Sequence, the model uses the same document to predict each target token. The second approach, RAG-Token, can predict each target token based on a different document. In the following, we formally introduce both models and then describe the p\u03b7 and p\u03b8 components, as well as the training and decoding procedure.", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2328 }, { "text": "2.1 Models RAG-Sequence Model The RAG-Sequence model uses the same retrieved document to generate the complete sequence. Technically, it treats the retrieved document as a single latent variable that is marginalized to get the seq2seq probability p(y|x) via a top-K approximation. Concretely, the top K documents are retrieved using the retriever, and the generator produces the output sequence", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2329 }, { "text": "probability for each document, which are then marginalized, pRAG-Sequence(y|x) \u2248 X z\u2208top-k(p(\u00b7|x)) p\u03b7(z|x)p\u03b8(y|x, z) = X z\u2208top-k(p(\u00b7|x)) p\u03b7(z|x) N Y i p\u03b8(yi|x, z, y1:i\u22121) RAG-Token Model In the RAG-Token model we can draw a different latent document for each target token and marginalize accordingly. This allows the generator to choose content from several", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2330 }, { "text": "documents when producing an answer. Concretely, the top K documents are retrieved using the retriever, and then the generator produces a distribution for the next output token for each document, before marginalizing, and repeating the process with the following output token, Formally, we de\ufb01ne: pRAG-Token(y|x) \u2248", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2331 }, { "text": "N Y i X z\u2208top-k(p(\u00b7|x)) p\u03b7(z|x)p\u03b8(yi|x, z, y1:i\u22121) Finally, we note that RAG can be used for sequence classi\ufb01cation tasks by considering the target class as a target sequence of length one, in which case RAG-Sequence and RAG-Token are equivalent. 2.2 Retriever: DPR The retrieval component p\u03b7(z|x) is based on DPR [26]. DPR follows a bi-encoder architecture:", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2332 }, { "text": "p\u03b7(z|x) \u221dexp \u0000d(z)\u22a4q(x) \u0001 d(z) = BERTd(z), q(x) = BERTq(x) where d(z) is a dense representation of a document produced by a BERTBASE document encoder [8], and q(x) a query representation produced by a query encoder, also based on BERTBASE. Calculating top-k(p\u03b7(\u00b7|x)), the list of k documents z with highest prior probability p\u03b7(z|x), is a Maximum Inner", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2333 }, { "text": "Product Search (MIPS) problem, which can be approximately solved in sub-linear time [23]. We use a pre-trained bi-encoder from DPR to initialize our retriever and to build the document index. This retriever was trained to retrieve documents which contain answers to TriviaQA [24] questions and Natural Questions [29]. We refer to the document index as the non-parametric memory.", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2334 }, { "text": "2.3 Generator: BART The generator component p\u03b8(yi|x, z, y1:i\u22121) could be modelled using any encoder-decoder. We use BART-large [32], a pre-trained seq2seq transformer [58] with 400M parameters. To combine the input x with the retrieved content z when generating from BART, we simply concatenate them. BART was", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2335 }, { "text": "pre-trained using a denoising objective and a variety of different noising functions. It has obtained state-of-the-art results on a diverse set of generation tasks and outperforms comparably-sized T5 models [32]. We refer to the BART generator parameters \u03b8 as the parametric memory henceforth. 2.4 Training", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2336 }, { "text": "minimize the negative marginal log-likelihood of each target, P j \u2212log p(yj|xj) using stochastic gradient descent with Adam [28]. Updating the document encoder BERTd during training is costly as it requires the document index to be periodically updated as REALM does during pre-training [20]. We do not \ufb01nd this step necessary for strong performance, and keep the document encoder (and", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2337 }, { "text": "index) \ufb01xed, only \ufb01ne-tuning the query encoder BERTq and the BART generator. 2.5 Decoding At test time, RAG-Sequence and RAG-Token require different ways to approximate arg maxy p(y|x). RAG-Token The RAG-Token model can be seen as a standard, autoregressive seq2seq genera- tor with transition probability: p\u2032", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2338 }, { "text": "\u03b8(yi|x, y1:i\u22121) = P z\u2208top-k(p(\u00b7|x)) p\u03b7(zi|x)p\u03b8(yi|x, zi, y1:i\u22121) To decode, we can plug p\u2032 \u03b8(yi|x, y1:i\u22121) into a standard beam decoder. RAG-Sequence For RAG-Sequence, the likelihood p(y|x) does not break into a conventional per- token likelihood, hence we cannot solve it with a single beam search. Instead, we run beam search for", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2339 }, { "text": "each document z, scoring each hypothesis using p\u03b8(yi|x, z, y1:i\u22121). This yields a set of hypotheses Y , some of which may not have appeared in the beams of all documents. To estimate the probability of an hypothesis y we run an additional forward pass for each document z for which y does not appear in the beam, multiply generator probability with p\u03b7(z|x) and then sum the probabilities across", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2340 }, { "text": "beams for the marginals. We refer to this decoding procedure as \u201cThorough Decoding.\u201d For longer output sequences, |Y | can become large, requiring many forward passes. For more ef\ufb01cient decoding, we can make a further approximation that p\u03b8(y|x, zi) \u22480 where y was not generated during beam search from x, zi. This avoids the need to run additional forward passes once the candidate set Y has", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2341 }, { "text": "been generated. We refer to this decoding procedure as \u201cFast Decoding.\u201d 3 Experiments We experiment with RAG in a wide range of knowledge-intensive tasks. For all experiments, we use a single Wikipedia dump for our non-parametric knowledge source. Following Lee et al. [31] and Karpukhin et al. [26], we use the December 2018 dump. Each Wikipedia article is split into disjoint", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2342 }, { "text": "100-word chunks, to make a total of 21M documents. We use the document encoder to compute an embedding for each document, and build a single MIPS index using FAISS [23] with a Hierarchical Navigable Small World approximation for fast retrieval [37]. During training, we retrieve the top k documents for each query. We consider k \u2208{5, 10} for training and set k for test time using dev", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2343 }, { "text": "data. We now discuss experimental details for each task. 3.1 Open-domain Question Answering Open-domain question answering (QA) is an important real-world application and common testbed for knowledge-intensive tasks [20]. We treat questions and answers as input-output text pairs (x, y) and train RAG by directly minimizing the negative log-likelihood of answers. We compare RAG to", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2344 }, { "text": "the popular extractive QA paradigm [5, 7, 31, 26], where answers are extracted spans from retrieved documents, relying primarily on non-parametric knowledge. We also compare to \u201cClosed-Book QA\u201d approaches [52], which, like RAG, generate answers, but which do not exploit retrieval, instead relying purely on parametric knowledge. We consider four popular open-domain QA datasets: Natural", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2345 }, { "text": "Questions (NQ) [29], TriviaQA (TQA) [24]. WebQuestions (WQ) [3] and CuratedTrec (CT) [2]. As CT and WQ are small, we follow DPR [26] by initializing CT and WQ models with our NQ RAG model. We use the same train/dev/test splits as prior work [31, 26] and report Exact Match (EM) scores. For TQA, to compare with T5 [52], we also evaluate on the TQA Wiki test set.", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2346 }, { "text": "3.2 Abstractive Question Answering RAG models can go beyond simple extractive QA and answer questions with free-form, abstractive text generation. To test RAG\u2019s natural language generation (NLG) in a knowledge-intensive setting, we use the MSMARCO NLG task v2.1 [43]. The task consists of questions, ten gold passages", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2347 }, { "text": "MSMARCO as an open-domain abstractive QA task. MSMARCO has some questions that cannot be answered in a way that matches the reference answer without access to the gold passages, such as \u201cWhat is the weather in Volcano, CA?\u201d so performance will be lower without using gold passages. We also note that some MSMARCO questions cannot be answered using Wikipedia alone. Here,", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2348 }, { "text": "RAG can rely on parametric knowledge to generate reasonable responses. 3.3 Jeopardy Question Generation To evaluate RAG\u2019s generation abilities in a non-QA setting, we study open-domain question gen- eration. Rather than use questions from standard open-domain QA tasks, which typically consist of short, simple questions, we propose the more demanding task of generating Jeopardy questions.", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2349 }, { "text": "Jeopardy is an unusual format that consists of trying to guess an entity from a fact about that entity. For example, \u201cThe World Cup\u201d is the answer to the question \u201cIn 1986 Mexico scored as the \ufb01rst country to host this international sports competition twice.\u201d As Jeopardy questions are precise, factual statements, generating Jeopardy questions conditioned on their answer entities constitutes a", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2350 }, { "text": "challenging knowledge-intensive generation task. We use the splits from SearchQA [10], with 100K train, 14K dev, and 27K test examples. As this is a new task, we train a BART model for comparison. Following [67], we evaluate using the SQuAD-tuned Q-BLEU-1 metric [42]. Q-BLEU is a variant of BLEU with a higher weight for", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2351 }, { "text": "matching entities and has higher correlation with human judgment for question generation than standard metrics. We also perform two human evaluations, one to assess generation factuality, and one for speci\ufb01city. We de\ufb01ne factuality as whether a statement can be corroborated by trusted external sources, and speci\ufb01city as high mutual dependence between the input and output [33]. We follow", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2352 }, { "text": "best practice and use pairwise comparative evaluation [34]. Evaluators are shown an answer and two generated questions, one from BART and one from RAG. They are then asked to pick one of four options\u2014quuestion A is better, question B is better, both are good, or neither is good. 3.4 Fact Veri\ufb01cation", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2353 }, { "text": "FEVER [56] requires classifying whether a natural language claim is supported or refuted by Wikipedia, or whether there is not enough information to decide. The task requires retrieving evidence from Wikipedia relating to the claim and then reasoning over this evidence to classify whether the claim is true, false, or unveri\ufb01able from Wikipedia alone. FEVER is a retrieval problem", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2354 }, { "text": "coupled with an challenging entailment reasoning task. It also provides an appropriate testbed for exploring the RAG models\u2019 ability to handle classi\ufb01cation rather than generation. We map FEVER class labels (supports, refutes, or not enough info) to single output tokens and directly train with claim-class pairs. Crucially, unlike most other approaches to FEVER, we do not use supervision on", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2355 }, { "text": "retrieved evidence. In many real-world applications, retrieval supervision signals aren\u2019t available, and models that do not require such supervision will be applicable to a wider range of tasks. We explore two variants: the standard 3-way classi\ufb01cation task (supports/refutes/not enough info) and the 2-way", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2356 }, { "text": "(supports/refutes) task studied in Thorne and Vlachos [57]. In both cases we report label accuracy. 4 Results 4.1 Open-domain Question Answering Table 1 shows results for RAG along with state-of-the-art models. On all four open-domain QA tasks, RAG sets a new state of the art (only on the T5-comparable split for TQA). RAG combines", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2357 }, { "text": "the generation \ufb02exibility of the \u201cclosed-book\u201d (parametric only) approaches and the performance of \"open-book\" retrieval-based approaches. Unlike REALM and T5+SSM, RAG enjoys strong results without expensive, specialized \u201csalient span masking\u201d pre-training [20]. It is worth noting that RAG\u2019s retriever is initialized using DPR\u2019s retriever, which uses retrieval supervision on Natural Questions", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2358 }, { "text": "and TriviaQA. RAG compares favourably to the DPR QA system, which uses a BERT-based \u201ccross- encoder\u201d to re-rank documents, along with an extractive reader. RAG demonstrates that neither a re-ranker nor extractive reader is necessary for state-of-the-art performance. There are several advantages to generating answers even when it is possible to extract them. Docu-", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2359 }, { "text": "40.4 - / - 40.7 46.8 DPR [26] 41.5 57.9/ - 41.1 50.6 RAG-Token 44.1 55.2/66.1 45.5 50.0 RAG-Seq. 44.5 56.8/68.0 45.2 52.2 Table 2: Generation and classi\ufb01cation Test Scores. MS-MARCO SotA is [4], FEVER-3 is [68] and FEVER-2 is [57] *Uses gold context/evidence. Best model without gold access underlined.", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2360 }, { "text": "Model Jeopardy MSMARCO FVR3 FVR2 B-1 QB-1 R-L B-1 Label Acc. SotA - - 49.8* 49.9* 76.8 92.2* BART 15.1 19.7 38.2 41.6 64.0 81.1 RAG-Tok. 17.3 22.2 40.1 41.5 72.5 89.5 RAG-Seq. 14.7 21.4 40.8 44.2 to more effective marginalization over documents. Furthermore, RAG can generate correct answers even when the correct answer is not in any retrieved document, achieving 11.8% accuracy in such", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2361 }, { "text": "cases for NQ, where an extractive model would score 0%. 4.2 Abstractive Question Answering As shown in Table 2, RAG-Sequence outperforms BART on Open MS-MARCO NLG by 2.6 Bleu points and 2.6 Rouge-L points. RAG approaches state-of-the-art model performance, which is impressive given that (i) those models access gold passages with speci\ufb01c information required to", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2362 }, { "text": "generate the reference answer , (ii) many questions are unanswerable without the gold passages, and (iii) not all questions are answerable from Wikipedia alone. Table 3 shows some generated answers from our models. Qualitatively, we \ufb01nd that RAG models hallucinate less and generate factually correct text more often than BART. Later, we also show that RAG generations are more diverse than", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2363 }, { "text": "BART generations (see \u00a74.5). 4.3 Jeopardy Question Generation Table 2 shows that RAG-Token performs better than RAG-Sequence on Jeopardy question generation, with both models outperforming BART on Q-BLEU-1. 4 shows human evaluation results, over 452 pairs of generations from BART and RAG-Token. Evaluators indicated that BART was more factual", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2364 }, { "text": "than RAG in only 7.1% of cases, while RAG was more factual in 42.7% of cases, and both RAG and BART were factual in a further 17% of cases, clearly demonstrating the effectiveness of RAG on the task over a state-of-the-art generation model. Evaluators also \ufb01nd RAG generations to be more speci\ufb01c by a large margin. Table 3 shows typical generations from each model.", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2365 }, { "text": "Jeopardy questions often contain two separate pieces of information, and RAG-Token may perform best because it can generate responses that combine content from several documents. Figure 2 shows an example. When generating \u201cSun\u201d, the posterior is high for document 2 which mentions \u201cThe Sun Also Rises\u201d. Similarly, document 1 dominates the posterior when \u201cA Farewell to Arms\u201d is", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2366 }, { "text": "generated. Intriguingly, after the \ufb01rst token of each book is generated, the document posterior \ufb02attens. This observation suggests that the generator can complete the titles without depending on speci\ufb01c documents. In other words, the model\u2019s parametric knowledge is suf\ufb01cient to complete the titles. We", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2367 }, { "text": "\ufb01nd evidence for this hypothesis by feeding the BART-only baseline with the partial decoding \"The Sun. BART completes the generation \"The Sun Also Rises\" is a novel by this author of \"The Sun Also Rises\" indicating the title \"The Sun Also Rises\" is stored in BART\u2019s parameters. Similarly, BART will complete the partial decoding \"The Sun Also Rises\" is a novel by this author of \"A", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2368 }, { "text": "with \"The Sun Also Rises\" is a novel by this author of \"A Farewell to Arms\". This example shows how parametric and non-parametric memories work together\u2014the non-parametric component helps to guide the generation, drawing out speci\ufb01c knowledge stored in the parametric memory. 4.4 Fact Veri\ufb01cation Table 2 shows our results on FEVER. For 3-way classi\ufb01cation, RAG scores are within 4.3% of", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2369 }, { "text": "Document 1: his works are considered classics of American literature ... His wartime experiences formed the basis for his novel \u201dA Farewell to Arms\u201d (1929) ... Document 2: ... artists of the 1920s \u201dLost Generation\u201d expatriate community. His debut novel, \u201dThe Sun Also Rises\u201d, was published in 1926. BOS \u201d", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2370 }, { "text": "The Sun Also R ises \u201d is a novel by this authorof \u201d A Fare well to Arms\u201d Doc 1 Doc 2 Doc 3 Doc 4 Doc 5 Figure 2: RAG-Token document posterior p(zi|x, yi, y\u2212i) for each generated token for input \u201cHem- ingway\" for Jeopardy generation with 5 retrieved documents. The posterior for document 1 is high when generating \u201cA Farewell to Arms\" and for document 2 when generating \u201cThe Sun Also Rises\".", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2371 }, { "text": "RAG-T The middle ear is the portion of the ear internal to the eardrum. RAG-S The middle ear includes the tympanic cavity and the three ossicles. what currency needed in scotland BART The currency needed in Scotland is Pound sterling. RAG-T Pound is the currency needed in Scotland. RAG-S The currency needed in Scotland is the pound sterling.", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2372 }, { "text": "Jeopardy Question Gener -ation Washington BART ?This state has the largest number of counties in the U.S. RAG-T It\u2019s the only U.S. state named for a U.S. president RAG-S It\u2019s the state where you\u2019ll \ufb01nd Mount Rainier National Park The Divine Comedy BART *This epic poem by Dante is divided into 3 parts: the Inferno, the Purgatorio & the Purgatorio", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2373 }, { "text": "RAG-T Dante\u2019s \"Inferno\" is the \ufb01rst part of this epic poem RAG-S This 14th century work is divided into 3 sections: \"Inferno\", \"Purgatorio\" & \"Paradiso\" For 2-way classi\ufb01cation, we compare against Thorne and Vlachos [57], who train RoBERTa [35] to classify the claim as true or false given the gold evidence sentence. RAG achieves an accuracy", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2374 }, { "text": "within 2.7% of this model, despite being supplied with only the claim and retrieving its own evidence. We also analyze whether documents retrieved by RAG correspond to documents annotated as gold evidence in FEVER. We calculate the overlap in article titles between the top k documents retrieved by RAG and gold evidence annotations. We \ufb01nd that the top retrieved document is from a gold article", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2375 }, { "text": "in 71% of cases, and a gold article is present in the top 10 retrieved articles in 90% of cases. 4.5 Additional Results Generation Diversity Section 4.3 shows that RAG models are more factual and speci\ufb01c than BART for Jeopardy question generation. Following recent work on diversity-promoting decoding", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2376 }, { "text": "[33, 59, 39], we also investigate generation diversity by calculating the ratio of distinct ngrams to total ngrams generated by different models. Table 5 shows that RAG-Sequence\u2019s generations are more diverse than RAG-Token\u2019s, and both are signi\ufb01cantly more diverse than BART without needing any diversity-promoting decoding.", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2377 }, { "text": "Retrieval Ablations A key feature of RAG is learning to retrieve relevant information for the task. To assess the effectiveness of the retrieval mechanism, we run ablations where we freeze the retriever during training. As shown in Table 6, learned retrieval improves results for all tasks. We compare RAG\u2019s dense retriever to a word overlap-based BM25 retriever [53]. Here, we replace", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2378 }, { "text": "RAG\u2019s retriever with a \ufb01xed BM25 system, and use BM25 retrieval scores as logits when calculating p(z|x). Table 6 shows the results. For FEVER, BM25 performs best, perhaps since FEVER claims are heavily entity-centric and thus well-suited for word overlap-based retrieval. Differentiable retrieval improves results on all other tasks, especially for Open-Domain QA, where it is crucial.", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2379 }, { "text": "Index hot-swapping An advantage of non-parametric memory models like RAG is that knowledge can be easily updated at test time. Parametric-only models like T5 or BART need further training to update their behavior as the world changes. To demonstrate, we build an index using the DrQA [5] Wikipedia dump from December 2016 and compare outputs from RAG using this index to the newer", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2380 }, { "text": "89.6% 90.0% BART 70.7% 32.4% RAG-Token 77.8% 46.8% RAG-Seq. 83.5% 53.8% Table 6: Ablations on the dev set. As FEVER is a classi\ufb01cation task, both RAG models are equivalent. Model NQ TQA WQ CT Jeopardy-QGen MSMarco FVR-3 FVR-2 Exact Match B-1 QB-1 R-L B-1 Label Accuracy RAG-Token-BM25 29.7 41.5 32.1", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2381 }, { "text": "33.1 17.5 22.3 55.5 48.4 75.1 91.6 RAG-Sequence-BM25 31.8 44.1 36.6 33.8 11.1 19.5 56.5 46.9 RAG-Token-Frozen 37.8 50.1 37.1 51.1 16.7 21.7 55.9 49.4 72.9 89.4 RAG-Sequence-Frozen 41.2 52.1 41.8 52.6 11.8 19.6 56.7 47.3 RAG-Token 43.5 54.8 46.5 51.9 17.9 22.6 56.2 49.4 74.5 90.6 RAG-Sequence 44.0 55.8", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2382 }, { "text": "44.9 53.4 15.3 21.5 57.2 47.5 between these dates and use a template \u201cWho is {position}?\u201d (e.g. \u201cWho is the President of Peru?\u201d) to query our NQ RAG model with each index. RAG answers 70% correctly using the 2016 index for 2016 world leaders and 68% using the 2018 index for 2018 world leaders. Accuracy with mismatched", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2383 }, { "text": "indices is low (12% with the 2018 index and 2016 leaders, 4% with the 2016 index and 2018 leaders). This shows we can update RAG\u2019s world knowledge by simply replacing its non-parametric memory. Effect of Retrieving more documents Models are trained with either 5 or 10 retrieved latent documents, and we do not observe signi\ufb01cant differences in performance between them. We have the", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2384 }, { "text": "\ufb02exibility to adjust the number of retrieved documents at test time, which can affect performance and runtime. Figure 3 (left) shows that retrieving more documents at test time monotonically improves Open-domain QA results for RAG-Sequence, but performance peaks for RAG-Token at 10 retrieved documents. Figure 3 (right) shows that retrieving more documents leads to higher Rouge-L for", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2385 }, { "text": "RAG-Token at the expense of Bleu-1, but the effect is less pronounced for RAG-Sequence. 10 20 30 40 50 K Retrieved Docs 39 40 41 42 43 44 NQ Exact Match RAG-Tok RAG-Seq 10 20 30 40 50 K Retrieved Docs 40 50 60 70 80 NQ Answer Recall @ K RAG-Tok RAG-Seq Fixed DPR BM25 10 20 30 40 50 K Retrieved Docs", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2386 }, { "text": "48 50 52 54 56 Bleu-1 / Rouge-L score RAG-Tok R-L RAG-Tok B-1 RAG-Seq R-L RAG-Seq B-1 Figure 3: Left: NQ performance as more documents are retrieved. Center: Retrieval recall perfor- mance in NQ. Right: MS-MARCO Bleu-1 and Rouge-L as more documents are retrieved. 5 Related Work Single-Task Retrieval", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2387 }, { "text": "Prior work has shown that retrieval improves performance across a variety of NLP tasks when considered in isolation. Such tasks include open-domain question answering [5, 29], fact checking [56], fact completion [48], long-form question answering [12], Wikipedia article generation [36], dialogue [41, 65, 9, 13], translation [17], and language modeling [19, 27]. Our", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2388 }, { "text": "General-Purpose Architectures for NLP Prior work on general-purpose architectures for NLP tasks has shown great success without the use of retrieval. A single, pre-trained language model has been shown to achieve strong performance on various classi\ufb01cation tasks in the GLUE bench- marks [60, 61] after \ufb01ne-tuning [49, 8]. GPT-2 [50] later showed that a single, left-to-right, pre-trained", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2389 }, { "text": "language model could achieve strong performance across both discriminative and generative tasks. For further improvement, BART [32] and T5 [51, 52] propose a single, pre-trained encoder-decoder model that leverages bi-directional attention to achieve stronger performance on discriminative and generative tasks. Our work aims to expand the space of possible tasks with a single, uni\ufb01ed", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2390 }, { "text": "architecture, by learning a retrieval module to augment pre-trained, generative language models. Learned Retrieval There is signi\ufb01cant work on learning to retrieve documents in information retrieval, more recently with pre-trained, neural language models [44, 26] similar to ours. Some work optimizes the retrieval module to aid in a speci\ufb01c, downstream task such as question answering,", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2391 }, { "text": "using search [46], reinforcement learning [6, 63, 62], or a latent variable approach [31, 20] as in our work. These successes leverage different retrieval-based architectures and optimization techniques to achieve strong performance on a single task, while we show that a single retrieval-based architecture", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2392 }, { "text": "can be \ufb01ne-tuned for strong performance on a variety of tasks. Memory-based Architectures Our document index can be seen as a large external memory for neural networks to attend to, analogous to memory networks [64, 55]. Concurrent work [14] learns to retrieve a trained embedding for each entity in the input, rather than to retrieve raw text as in our", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2393 }, { "text": "work. Other work improves the ability of dialog models to generate factual text by attending over fact embeddings [15, 13]. A key feature of our memory is that it is comprised of raw text rather distributed representations, which makes the memory both (i) human-readable, lending a form of interpretability to our model, and (ii) human-writable, enabling us to dynamically update the model\u2019s", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2394 }, { "text": "memory by editing the document index. This approach has also been used in knowledge-intensive dialog, where generators have been conditioned on retrieved text directly, albeit obtained via TF-IDF rather than end-to-end learnt retrieval [9]. Retrieve-and-Edit approaches Our method shares some similarities with retrieve-and-edit style", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2395 }, { "text": "approaches, where a similar training input-output pair is retrieved for a given input, and then edited to provide a \ufb01nal output. These approaches have proved successful in a number of domains including Machine Translation [18, 22] and Semantic Parsing [21]. Our approach does have several differences,", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2396 }, { "text": "including less of emphasis on lightly editing a retrieved item, but on aggregating content from several pieces of retrieved content, as well as learning latent retrieval, and retrieving evidence documents rather than related training pairs. This said, RAG techniques may work well in these settings, and", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2397 }, { "text": "could represent promising future work. 6 Discussion In this work, we presented hybrid generation models with access to parametric and non-parametric memory. We showed that our RAG models obtain state of the art results on open-domain QA. We found that people prefer RAG\u2019s generation over purely parametric BART, \ufb01nding RAG more factual", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2398 }, { "text": "and speci\ufb01c. We conducted an thorough investigation of the learned retrieval component, validating its effectiveness, and we illustrated how the retrieval index can be hot-swapped to update the model without requiring any retraining. In future work, it may be fruitful to investigate if the two components", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2399 }, { "text": "can be jointly pre-trained from scratch, either with a denoising objective similar to BART or some another objective. Our work opens up new research directions on how parametric and non-parametric memories interact and how to most effectively combine them, showing promise in being applied to a wide variety of NLP tasks.", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2400 }, { "text": "Broader Impact This work offers several positive societal bene\ufb01ts over previous work: the fact that it is more strongly grounded in real factual knowledge (in this case Wikipedia) makes it \u201challucinate\u201d less with generations that are more factual, and offers more control and interpretability. RAG could be", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2401 }, { "text": "employed in a wide variety of scenarios with direct bene\ufb01t to society, for example by endowing it with a medical index and asking it open-domain questions on that topic, or by helping people be more effective at their jobs. With these advantages also come potential downsides: Wikipedia, or any potential external knowledge", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2402 }, { "text": "source, will probably never be entirely factual and completely devoid of bias. Since RAG can be employed as a language model, similar concerns as for GPT-2 [50] are valid here, although arguably to a lesser extent, including that it might be used to generate abuse, faked or misleading content in the news or on social media; to impersonate others; or to automate the production of spam/phishing", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2403 }, { "text": "content [54]. Advanced language models may also lead to the automation of various jobs in the coming decades [16]. In order to mitigate these risks, AI systems could be employed to \ufb01ght against misleading content and automated spam/phishing. Acknowledgments The authors would like to thank the reviewers for their thoughtful and constructive feedback on this", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2404 }, { "text": "paper, as well as HuggingFace for their help in open-sourcing code to run RAG models. The authors would also like to thank Kyunghyun Cho and Sewon Min for productive discussions and advice. EP thanks supports from the NSF Graduate Research Fellowship. PL is supported by the FAIR PhD program. References", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2405 }, { "text": "//arxiv.org/abs/1611.09268. arXiv: 1611.09268. [2] Petr Baudi\u0161 and Jan \u0160ediv`y. Modeling of the question answering task in the yodaqa system. In International Conference of the Cross-Language Evaluation Forum for European Languages, pages 222\u2013228. Springer, 2015. URL https://link.springer.com/chapter/10.1007%", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2406 }, { "text": "2F978-3-319-24027-5_20. [3] Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. Semantic Parsing on Freebase from Question-Answer Pairs. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pages 1533\u20131544, Seattle, Washington, USA, October 2013. Association for Computational Linguistics. URL http://www.aclweb.org/anthology/", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2407 }, { "text": "D13-1160. [4] Bin Bi, Chenliang Li, Chen Wu, Ming Yan, and Wei Wang. Palm: Pre-training an autoencod- ing&autoregressive language model for context-conditioned generation. ArXiv, abs/2004.07159, 2020. URL https://arxiv.org/abs/2004.07159. [5] Danqi Chen, Adam Fisch, Jason Weston, and Antoine Bordes. Reading Wikipedia to Answer", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2408 }, { "text": "Open-Domain Questions. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1870\u20131879, Vancouver, Canada, July 2017. Association for Computational Linguistics. doi: 10.18653/v1/P17-1171. URL https://www.aclweb.org/anthology/P17-1171.", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2409 }, { "text": "ference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171\u20134186, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1423. URL https://www.aclweb.org/anthology/N19-1423.", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2410 }, { "text": "Cho. SearchQA: A New Q&A Dataset Augmented with Context from a Search Engine. arXiv:1704.05179 [cs], April 2017. URL http://arxiv.org/abs/1704.05179. arXiv: 1704.05179. [11] Angela Fan, Mike Lewis, and Yann Dauphin. Hierarchical neural story generation. In Proceed- ings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1:", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2411 }, { "text": "Long Papers), pages 889\u2013898, Melbourne, Australia, July 2018. Association for Computational Linguistics. doi: 10.18653/v1/P18-1082. URL https://www.aclweb.org/anthology/ P18-1082. [12] Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. ELI5: Long form question answering. In Proceedings of the 57th Annual Meeting of the Association", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2412 }, { "text": "for Computational Linguistics, pages 3558\u20133567, Florence, Italy, July 2019. Association for Computational Linguistics. doi: 10.18653/v1/P19-1346. URL https://www.aclweb.org/ anthology/P19-1346. [13] Angela Fan, Claire Gardent, Chloe Braud, and Antoine Bordes. Augmenting transformers with KNN-based composite memory, 2020. URL https://openreview.net/forum?id=", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2413 }, { "text": "H1gx1CNKPH. [14] Thibault F\u00e9vry, Livio Baldini Soares, Nicholas FitzGerald, Eunsol Choi, and Tom Kwiatkowski. Entities as experts: Sparse memory access with entity supervision. ArXiv, abs/2004.07202, 2020. URL https://arxiv.org/abs/2004.07202. [15] Marjan Ghazvininejad, Chris Brockett, Ming-Wei Chang, Bill Dolan, Jianfeng Gao, Wen", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2414 }, { "text": "tau Yih, and Michel Galley. A knowledge-grounded neural conversation model. In AAAI Conference on Arti\ufb01cial Intelligence, 2018. URL https://www.aaai.org/ocs/index.php/ AAAI/AAAI18/paper/view/16710. [16] Katja Grace, John Salvatier, Allan Dafoe, Baobao Zhang, and Owain Evans. When will AI exceed human performance? evidence from AI experts. CoRR, abs/1705.08807, 2017. URL", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2415 }, { "text": "http://arxiv.org/abs/1705.08807. [17] Jiatao Gu, Yong Wang, Kyunghyun Cho, and Victor O.K. Li. Search engine guided neural machine translation. In AAAI Conference on Arti\ufb01cial Intelligence, 2018. URL https: //www.aaai.org/ocs/index.php/AAAI/AAAI18/paper/view/17282. [18] Jiatao Gu, Yong Wang, Kyunghyun Cho, and Victor O.K. Li. Search engine guided neural", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2416 }, { "text": "machine translation. In 32nd AAAI Conference on Arti\ufb01cial Intelligence, AAAI 2018, 32nd AAAI Conference on Arti\ufb01cial Intelligence, AAAI 2018, pages 5133\u20135140. AAAI press, 2018. 32nd AAAI Conference on Arti\ufb01cial Intelligence, AAAI 2018 ; Conference date: 02-02-2018 Through 07-02-2018. [19] Kelvin Guu, Tatsunori B. Hashimoto, Yonatan Oren, and Percy Liang. Generating sentences by", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2417 }, { "text": "for predicting structured outputs. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, ed- itors, Advances in Neural Information Processing Systems 31, pages 10052\u2013 10062. Curran Associates, Inc., 2018. URL http://papers.nips.cc/paper/ 8209-a-retrieve-and-edit-framework-for-predicting-structured-outputs.", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2418 }, { "text": "pdf. [22] Nabil Hossain, Marjan Ghazvininejad, and Luke Zettlemoyer. Simple and effective retrieve- edit-rerank text generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2532\u20132538, Online, July 2020. Association for Computa- tional Linguistics. doi: 10.18653/v1/2020.acl-main.228. URL https://www.aclweb.org/", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2419 }, { "text": "anthology/2020.acl-main.228. [23] Jeff Johnson, Matthijs Douze, and Herv\u00e9 J\u00e9gou. Billion-scale similarity search with gpus. arXiv preprint arXiv:1702.08734, 2017. URL https://arxiv.org/abs/1702.08734. [24] Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension. In Proceedings of the", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2420 }, { "text": "55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1601\u20131611, Vancouver, Canada, July 2017. Association for Computational Linguistics. doi: 10.18653/v1/P17-1147. URL https://www.aclweb.org/anthology/P17-1147. [25] Armand Joulin and Tomas Mikolov. Inferring algorithmic patterns with stack-", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2421 }, { "text": "augmented recurrent nets. In Proceedings of the 28th International Conference on Neural Information Processing Systems - Volume 1, NIPS\u201915, page 190\u2013198, Cam- bridge, MA, USA, 2015. MIT Press. URL https://papers.nips.cc/paper/ 5857-inferring-algorithmic-patterns-with-stack-augmented-recurrent-nets.", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2422 }, { "text": "tion through memorization: Nearest neighbor language models. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=HklBjCEKvH. [28] Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In Yoshua Bengio and Yann LeCun, editors, 3rd International Conference on Learning Representations,", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2423 }, { "text": "ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, 2015. URL http://arxiv.org/abs/1412.6980. [29] Tom Kwiatkowski, Jennimaria Palomaki, Olivia Red\ufb01eld, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Matthew Kelcey, Jacob Devlin, Ken- ton Lee, Kristina N. Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2424 }, { "text": "Uszkoreit, Quoc Le, and Slav Petrov. Natural Questions: a Benchmark for Ques- tion Answering Research. Transactions of the Association of Computational Lin- guistics, 2019. URL https://tomkwiat.users.x20web.corp.google.com/papers/ natural-questions/main-1455-kwiatkowski.pdf. [30] Guillaume Lample, Alexandre Sablayrolles, Marc\u2019 Aurelio Ranzato, Ludovic Denoyer, and", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2425 }, { "text": "Herve Jegou. Large memory layers with product keys. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d\u2019 Alch\u00e9-Buc, E. Fox, and R. Garnett, editors, Advances in Neural In- formation Processing Systems 32, pages 8548\u20138559. Curran Associates, Inc., 2019. URL http: //papers.nips.cc/paper/9061-large-memory-layers-with-product-keys.pdf.", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2426 }, { "text": "for Computational Linguistics, pages 6086\u20136096, Florence, Italy, July 2019. Association for Computational Linguistics. doi: 10.18653/v1/P19-1612. URL https://www.aclweb.org/ anthology/P19-1612. [32] Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. BART: Denoising sequence-to-sequence", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2427 }, { "text": "pre-training for natural language generation, translation, and comprehension. arXiv preprint arXiv:1910.13461, 2019. URL https://arxiv.org/abs/1910.13461. [33] Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. A diversity-promoting objective function for neural conversation models. In Proceedings of the 2016 Conference of the", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2428 }, { "text": "North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 110\u2013119, San Diego, California, June 2016. Association for Computational Linguistics. doi: 10.18653/v1/N16-1014. URL https://www.aclweb.org/anthology/ N16-1014. [34] Margaret Li, Jason Weston, and Stephen Roller. Acute-eval: Improved dialogue evaluation", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2429 }, { "text": "with optimized questions and multi-turn comparisons. ArXiv, abs/1909.03087, 2019. URL https://arxiv.org/abs/1909.03087. [35] Hairong Liu, Mingbo Ma, Liang Huang, Hao Xiong, and Zhongjun He. Robust neural machine translation with joint textual and phonetic embedding. In Proceedings of the 57th Annual", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2430 }, { "text": "Meeting of the Association for Computational Linguistics, pages 3044\u20133049, Florence, Italy, July 2019. Association for Computational Linguistics. doi: 10.18653/v1/P19-1291. URL https://www.aclweb.org/anthology/P19-1291. [36] Peter J. Liu*, Mohammad Saleh*, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser,", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2431 }, { "text": "and Noam Shazeer. Generating wikipedia by summarizing long sequences. In International Conference on Learning Representations, 2018. URL https://openreview.net/forum? id=Hyg0vbWC-. [37] Yury A. Malkov and D. A. Yashunin. Ef\ufb01cient and robust approximate nearest neighbor search using hierarchical navigable small world graphs. IEEE Transactions on Pattern Analysis and", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2432 }, { "text": "Machine Intelligence, 42:824\u2013836, 2016. URL https://arxiv.org/abs/1603.09320. [38] Gary Marcus. The next decade in ai: four steps towards robust arti\ufb01cial intelligence. arXiv preprint arXiv:2002.06177, 2020. URL https://arxiv.org/abs/2002.06177. [39] Luca Massarelli, Fabio Petroni, Aleksandra Piktus, Myle Ott, Tim Rockt\u00e4schel, Vassilis", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2433 }, { "text": "Plachouras, Fabrizio Silvestri, and Sebastian Riedel. How decoding strategies affect the veri\ufb01ability of generated text. arXiv preprint arXiv:1911.03587, 2019. URL https: //arxiv.org/abs/1911.03587. [40] Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu. Mixed", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2434 }, { "text": "precision training. In ICLR, 2018. URL https://openreview.net/forum?id=r1gs9JgRZ. [41] Nikita Moghe, Siddhartha Arora, Suman Banerjee, and Mitesh M. Khapra. Towards exploit- ing background knowledge for building conversation systems. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2322\u20132332, Brus-", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2435 }, { "text": "sels, Belgium, October-November 2018. Association for Computational Linguistics. doi: 10.18653/v1/D18-1255. URL https://www.aclweb.org/anthology/D18-1255. [42] Preksha Nema and Mitesh M. Khapra. Towards a better metric for evaluating question generation systems. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2436 }, { "text": "Processing, pages 3950\u20133959, Brussels, Belgium, October-November 2018. Association for Computational Linguistics. doi: 10.18653/v1/D18-1429. URL https://www.aclweb.org/ anthology/D18-1429. [43] Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. MS MARCO: A human generated machine reading comprehension dataset. In", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2437 }, { "text": "approaches 2016 co-located with the 30th Annual Conference on Neural Information Processing Systems (NIPS 2016), Barcelona, Spain, December 9, 2016, volume 1773 of CEUR Workshop Proceedings. CEUR-WS.org, 2016. URL http://ceur-ws.org/Vol-1773/CoCoNIPS_ 2016_paper9.pdf. [44] Rodrigo Nogueira and Kyunghyun Cho. Passage re-ranking with BERT. arXiv preprint", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2438 }, { "text": "arXiv:1901.04085, 2019. URL https://arxiv.org/abs/1901.04085. [45] Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. fairseq: A fast, extensible toolkit for sequence modeling. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2439 }, { "text": "Linguistics (Demonstrations), pages 48\u201353, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-4009. URL https://www.aclweb. org/anthology/N19-4009. [46] Ethan Perez, Siddharth Karamcheti, Rob Fergus, Jason Weston, Douwe Kiela, and Kyunghyun Cho. Finding generalizable evidence by learning to convince q&a models. In Proceedings", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2440 }, { "text": "of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2402\u20132411, Hong Kong, China, November 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1244. URL https://www.aclweb.org/anthology/D19-1244.", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2441 }, { "text": "Kong, China, November 2019. Association for Computational Linguistics. doi: 10.18653/v1/ D19-1250. URL https://www.aclweb.org/anthology/D19-1250. [48] Fabio Petroni, Patrick Lewis, Aleksandra Piktus, Tim Rockt\u00e4schel, Yuxiang Wu, Alexander H. Miller, and Sebastian Riedel. How context affects language models\u2019 factual predictions. In", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2442 }, { "text": "Automated Knowledge Base Construction, 2020. URL https://openreview.net/forum? id=025X0zPfn. [49] Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Im- proving Language Understanding by Generative Pre-Training, 2018. URL https://s3-us-west-2.amazonaws.com/openai-assets/research-covers/", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2443 }, { "text": "language-unsupervised/language_understanding_paper.pdf. [50] Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners, 2019. URL https://d4mucfpksywv.cloudfront.net/better-language-models/language_ models_are_unsupervised_multitask_learners.pdf.", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2444 }, { "text": "the parameters of a language model? arXiv e-prints, 2020. URL https://arxiv.org/abs/ 2002.08910. [53] Stephen Robertson and Hugo Zaragoza. The probabilistic relevance framework: Bm25 and beyond. Found. Trends Inf. Retr., 3(4):333\u2013389, April 2009. ISSN 1554-0669. doi: 10.1561/ 1500000019. URL https://doi.org/10.1561/1500000019.", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2445 }, { "text": "June 2018. Association for Computational Linguistics. doi: 10.18653/v1/N18-1074. URL https://www.aclweb.org/anthology/N18-1074. [57] James H. Thorne and Andreas Vlachos. Avoiding catastrophic forgetting in mitigating model biases in sentence-pair classi\ufb01cation with elastic weight consolidation. ArXiv, abs/2004.14366,", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2446 }, { "text": "2020. URL https://arxiv.org/abs/2004.14366. [58] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, \u0141 ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2447 }, { "text": "Information Processing Systems 30, pages 5998\u20136008. Curran Associates, Inc., 2017. URL http://papers.nips.cc/paper/7181-attention-is-all-you-need.pdf. [59] Ashwin Vijayakumar, Michael Cogswell, Ramprasaath Selvaraju, Qing Sun, Stefan Lee, David Crandall, and Dhruv Batra. Diverse beam search for improved description of complex scenes.", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2448 }, { "text": "AAAI Conference on Arti\ufb01cial Intelligence, 2018. URL https://www.aaai.org/ocs/index. php/AAAI/AAAI18/paper/view/17329. [60] Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2449 }, { "text": "Neural Networks for NLP, pages 353\u2013355, Brussels, Belgium, November 2018. Association for Computational Linguistics. doi: 10.18653/v1/W18-5446. URL https://www.aclweb.org/ anthology/W18-5446. [61] Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. SuperGLUE: A Stickier Benchmark for General-", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2450 }, { "text": "Purpose Language Understanding Systems. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d\\textquotesingle Alch\u00e9-Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems 32, pages 3261\u20133275. Curran Associates, Inc., 2019. URL https:// arxiv.org/abs/1905.00537. [62] Shuohang Wang, Mo Yu, Xiaoxiao Guo, Zhiguo Wang, Tim Klinger, Wei Zhang, Shiyu Chang,", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2451 }, { "text": "Gerry Tesauro, Bowen Zhou, and Jing Jiang. R3: Reinforced ranker-reader for open-domain question answering. In Sheila A. McIlraith and Kilian Q. Weinberger, editors, Proceedings of the Thirty-Second AAAI Conference on Arti\ufb01cial Intelligence, (AAAI-18), the 30th innovative Applications of Arti\ufb01cial Intelligence (IAAI-18), and the 8th AAAI Symposium on Educational", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2452 }, { "text": "Advances in Arti\ufb01cial Intelligence (EAAI-18), New Orleans, Louisiana, USA, February 2-7, 2018, pages 5981\u20135988. AAAI Press, 2018. URL https://www.aaai.org/ocs/index. php/AAAI/AAAI18/paper/view/16712. [63] Shuohang Wang, Mo Yu, Jing Jiang, Wei Zhang, Xiaoxiao Guo, Shiyu Chang, Zhiguo Wang, Tim Klinger, Gerald Tesauro, and Murray Campbell. Evidence aggregation for answer re-", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2453 }, { "text": "ranking in open-domain question answering. In ICLR, 2018. URL https://openreview. net/forum?id=rJl3yM-Ab. [64] Jason Weston, Sumit Chopra, and Antoine Bordes. Memory networks. In Yoshua Bengio and Yann LeCun, editors, 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, 2015. URL", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2454 }, { "text": "http://arxiv.org/abs/1410.3916. [65] Jason Weston, Emily Dinan, and Alexander Miller. Retrieve and re\ufb01ne: Improved sequence generation models for dialogue. In Proceedings of the 2018 EMNLP Workshop SCAI: The 2nd International Workshop on Search-Oriented Conversational AI, pages 87\u201392, Brussels, Belgium,", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2455 }, { "text": "State-of-the-art natural language processing. ArXiv, abs/1910.03771, 2019. [67] Shiyue Zhang and Mohit Bansal. Addressing semantic drift in question generation for semi- supervised question answering. In Proceedings of the 2019 Conference on Empirical Meth- ods in Natural Language Processing and the 9th International Joint Conference on Natural", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2456 }, { "text": "Language Processing (EMNLP-IJCNLP), pages 2495\u20132509, Hong Kong, China, Novem- ber 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1253. URL https://www.aclweb.org/anthology/D19-1253. [68] Wanjun Zhong, Jingjing Xu, Duyu Tang, Zenan Xu, Nan Duan, Ming Zhou, Jiahai Wang, and Jian Yin. Reasoning over semantic-level graph for fact checking. ArXiv, abs/1909.03745, 2019.", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2457 }, { "text": "Appendices for Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks A Implementation Details For Open-domain QA we report test numbers using 15 retrieved documents for RAG-Token models. For RAG-Sequence models, we report test results using 50 retrieved documents, and we use the Thorough Decoding approach since answers are generally short. We use greedy decoding for QA as", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2458 }, { "text": "we did not \ufb01nd beam search improved results. For Open-MSMarco and Jeopardy question generation, we report test numbers using ten retrieved documents for both RAG-Token and RAG-Sequence, and we also train a BART-large model as a baseline. We use a beam size of four, and use the Fast Decoding approach for RAG-Sequence models, as Thorough Decoding did not improve performance.", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2459 }, { "text": "B Human Evaluation Figure 4: Annotation interface for human evaluation of factuality. A pop-out for detailed instructions and a worked example appear when clicking \"view tool guide\". Figure 4 shows the user interface for human evaluation. To avoid any biases for screen position, which model corresponded to sentence A and sentence B was randomly selected for each example.", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2460 }, { "text": "Annotators were encouraged to research the topic using the internet, and were given detailed instruc- tions and worked examples in a full instructions tab. We included some gold sentences in order to assess the accuracy of the annotators. Two annotators did not perform well on these examples and their annotations were removed from the results.", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2461 }, { "text": "C Training setup Details We train all RAG models and BART baselines using Fairseq [45].2 We train with mixed precision \ufb02oating point arithmetic [40], distributing training across 8, 32GB NVIDIA V100 GPUs, though training and inference can be run on one GPU. We \ufb01nd that doing Maximum Inner Product Search", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2462 }, { "text": "with FAISS is suf\ufb01ciently fast on CPU, so we store document index vectors on CPU, requiring \u223c100 GB of CPU memory for all of Wikipedia. After submission, We have ported our code to HuggingFace Transformers [66]3, which achieves equivalent performance to the previous version but is a cleaner and easier to use implementation. This version is also open-sourced. We also compress the document", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2463 }, { "text": "index using FAISS\u2019s compression tools, reducing the CPU memory requirement to 36GB. Scripts to run experiments with RAG can be found at https://github.com/huggingface/transformers/ blob/master/examples/rag/README.md and an interactive demo of a RAG model can be found at https://huggingface.co/rag/ 2https://github.com/pytorch/fairseq", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2464 }, { "text": "D Further Details on Open-Domain QA For open-domain QA, multiple answer annotations are often available for a given question. These answer annotations are exploited by extractive models during training as typically all the answer annotations are used to \ufb01nd matches within documents when preparing training data. For RAG, we", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2465 }, { "text": "also make use of multiple annotation examples for Natural Questions and WebQuestions by training the model with each (q, a) pair separately, leading to a small increase in accuracy. For TriviaQA, there are often many valid answers to a given question, some of which are not suitable training targets,", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2466 }, { "text": "such as emoji or spelling variants. For TriviaQA, we \ufb01lter out answer candidates if they do not occur in top 1000 documents for the query. CuratedTrec preprocessing The answers for CuratedTrec are given in the form of regular expres- sions, which has been suggested as a reason why it is unsuitable for answer-generation models [20].", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2467 }, { "text": "To overcome this, we use a pre-processing step where we \ufb01rst retrieve the top 1000 documents for each query, and use the answer that most frequently matches the regex pattern as the supervision target. If no matches are found, we resort to a simple heuristic: generate all possible permutations for each regex, replacing non-deterministic symbols in the regex nested tree structure with a whitespace.", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2468 }, { "text": "TriviaQA Evaluation setups The open-domain QA community customarily uses public develop- ment datasets as test datasets, as test data for QA datasets is often restricted and dedicated to reading compehension purposes. We report our results using the datasets splits used in DPR [26], which are consistent with common practice in Open-domain QA. For TriviaQA, this test dataset is the public", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2469 }, { "text": "TriviaQA Web Development split. Roberts et al. [52] used the TriviaQA of\ufb01cial Wikipedia test set instead. F\u00e9vry et al. [14] follow this convention in order to compare with Roberts et al. [52] (See appendix of [14]). We report results on both test sets to enable fair comparison to both approaches. We \ufb01nd that our performance is much higher using the of\ufb01cial Wiki test set, rather than the more", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2470 }, { "text": "conventional open-domain test set, which we attribute to the of\ufb01cial Wiki test set questions being simpler to answer from Wikipedia. E Further Details on FEVER For FEVER classi\ufb01cation, we follow the practice from [32], and \ufb01rst re-generate the claim, and then classify using the representation of the \ufb01nal hidden state, before \ufb01nally marginalizing across", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2471 }, { "text": "documents to obtain the class probabilities. The FEVER task traditionally has two sub-tasks. The \ufb01rst is to classify the claim as either \"Supported\", \"Refuted\" or \"Not Enough Info\", which is the task we explore in the main paper. FEVER\u2019s other sub-task involves extracting sentences from Wikipedia as evidence supporting the classi\ufb01cation prediction. As FEVER uses a different Wikipedia dump to", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2472 }, { "text": "us, directly tackling this task is not straightforward. We hope to address this in future work. F Null Document Probabilities We experimented with adding \"Null document\" mechanism to RAG, similar to REALM [20] in order to model cases where no useful information could be retrieved for a given input. Here, if k documents", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2473 }, { "text": "were retrieved, we would additionally \"retrieve\" an empty document and predict a logit for the null document, before marginalizing over k + 1 predictions. We explored modelling this null document logit by learning (i) a document embedding for the null document, (ii) a static learnt bias term, or (iii) a neural network to predict the logit. We did not \ufb01nd that these improved performance, so in", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2474 }, { "text": "the interests of simplicity, we omit them. For Open MS-MARCO, where useful retrieved documents cannot always be retrieved, we observe that the model learns to always retrieve a particular set of documents for questions that are less likely to bene\ufb01t from retrieval, suggesting that null document mechanisms may not be necessary for RAG.", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2475 }, { "text": "12468 101093* FEVER-3-way 145450 10000 10000 FEVER-2-way 96966 6666 6666 parameters. The best performing \"closed-book\" (parametric only) open-domain QA model is T5-11B with 11 Billion trainable parameters. The T5 model with the closest number of parameters to our models is T5-large (770M parameters), which achieves a score of 28.9 EM on Natural Questions [52],", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2476 }, { "text": "substantially below the 44.5 that RAG-Sequence achieves, indicating that hybrid parametric/non- parametric models require far fewer trainable parameters for strong open-domain QA performance. The non-parametric memory index does not consist of trainable parameters, but does consists of 21M 728 dimensional vectors, consisting of 15.3B values. These can be easily be stored at 8-bit \ufb02oating", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2477 }, { "text": "point precision to manage memory and disk footprints. H Retrieval Collapse In preliminary experiments, we observed that for some tasks such as story generation [11], the retrieval component would \u201ccollapse\u201d and learn to retrieve the same documents regardless of the input. In these cases, once retrieval had collapsed, the generator would learn to ignore the documents,", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2478 }, { "text": "and the RAG model would perform equivalently to BART. The collapse could be due to a less-explicit requirement for factual knowledge in some tasks, or the longer target sequences, which could result in less informative gradients for the retriever. Perez et al. [46] also found spurious retrieval results", "source": "RAG Lewis et al", "year": 2020, "url": "https://arxiv.org/abs/2005.11401", "id": 2479 } ]