- 08/09/2025
- Category: Commentaries
Author: Fathin Difa Robbani
Editor: Achmed Faiz Yudha Siregar
The world of AI today has reached a level that’s both exciting and a bit terrifying. Imagine this: AI can now generate videos and images so realistic that it’s becoming nearly impossible for humans to tell what’s real and what’s fake, just like what you see in Figure 1.

Figure 1. Example of Veo3-Generated Video (Source: YouTube, 2024)
The spread of generative AI videos is no longer just a tech concern, it’s a real societal issue. Take the 2024 general election in India, for example. Deepfake videos were everywhere, to the point where even regular people couldn’t distinguish between authentic content and political manipulation (BBC News, 2024). According to BBC’s 2024 report, thousands of AI-generated videos went viral, misleading the public, making it harder to trust the news, and even stirring political tension. This shows just how urgently we need reliable technology to detect AI-generated content, especially to safeguard public discourse and critical events like elections from advanced digital deception.
With AI-generated content becoming more lifelike, many people are now asking: “How can we tell if a video is real or AI-made?” The potential for misuse is rising, hoaxes, deepfakes, and disinformation are becoming easier to create and spread.
If AI has become this good at “fooling” humans, then we need AI that’s equally, or even more advanced, to defend against it. In other words, we need a balance (AI versus AI) to keep technology safe for people.
In the AI research world, the topic of AI ethics is gaining momentum. It’s about ensuring AI remains beneficial rather than harmful to society. For instance, while Google developed Veo 3 (a powerful AI for video generation) they also focused on safety by building a digital watermarking system called SynthID (DeepMind, 2024). What makes SynthID remarkable is that its watermark is completely invisible to the human eye. Even after compressing, cropping, or editing the image, the watermark remains detectable by AI. That’s how smart it is.
Google’s goal with SynthID is simple but crucial: to maintain public trust amidst the rising tide of generative AI content. They also aim to comply with increasingly strict global regulations on AI and deepfakes. Unlike platforms that rely solely on bans or takedowns, Google chose a more forward-thinking route, embedding invisible watermarks in their AI-generated content. This ensures moderation can continue without ruining the user experience. Google has already been known for taking content moderation more seriously, like on YouTube and Google Search, and SynthID reinforces its position as a pioneer in transparent AI content governance (Google, 2023; YouTube Official Blog, 2024).
The Challenge of Detecting AI-Generated Video
In the past, we could spot AI-generated fakes with just our eyes. But now? Not anymore. The quality of AI-generated video has reached near-perfection. Old-school methods like visible watermarks can be easily removed or damaged. Think about how easily the DALL-E watermark can be erased or cropped. AI-made videos can be altered, resized, or remixed to the point that they’re indistinguishable from real footage.
So in this new era, where AI is “smart” enough to fool humans, it’s clear: our AI detection systems need to be even smarter. Just like in cybersecurity, offense and defense must evolve in tandem. That’s why AI ethics is such a hot topic in research, how do we ensure AI remains a force for good, not a double-edged sword?
How Does SynthID Work?
Referring to the development of invisible watermarking, this section will dive into some technical details, but don’t worry, it will be explained simply and with relatable analogies (Kohli, 2024; Zhao, 2024).
Let’s start with how AI understands digital images. As we know, digital images stored on a computer are essentially a collection of pixels. The higher the resolution (say, 1080p), the more pixels it contains, and usually, the sharper the image appears. Each pixel typically holds a value between 0–255 to represent color. For colored images, each pixel has three numbers: red, green, and blue (RGB).
But AI doesn’t “see” images pixel by pixel like humans do. Instead, it looks at groups of pixels at once using what’s called a filter. A filter is usually a small grid (like 3×3 pixels) that slides across the image to observe patterns. Think of it like using a magnifying glass to scan a painting, moving from left to right, top to bottom. That magnifying glass is what the filter acts like. The result of this sliding process is what the AI uses to interpret the image. This result is known as a feature map, an abstract representation of the image that only the AI understands.
This filtering process can be stacked layer by layer, just like how the human brain processes visual information in multiple layers. The deeper the layer, the more refined the AI’s understanding (although in reality, going too deep can sometimes make things worse rather than better).
To make the AI smarter, every time it “looks” at an image, it learns how to improve its understanding through a process called loss and backpropagation. In simple terms, the AI gets “rewarded” or “punished” based on how close its interpretation is to the ground truth, and then it adjusts its internal parameters over and over again to get better.
Hiding a Secret Code Inside an AI Image
Here’s where SynthID gets clever: the watermark isn’t embedded directly into visible pixels, but hidden inside the feature map, a representation of the image that’s already been abstracted into numbers. The value inserted is extremely small, usually below zero, so it doesn’t affect the image’s visual appearance at all (Zhao, 2024).
The process begins after the AI-generated image is completed. That image is then passed into a special AI model (which also uses filters), producing the feature map we mentioned earlier. At the same time, a secret code (say, 8-bit or 16-bit) is encoded, converted into a numerical “map” that has the same dimensions as the feature map (Kohli, 2024).
Then, the information from the secret code and the feature map is combined, perhaps via addition, concatenation, or other neural network techniques. Because the code’s values are so small, the resulting changes in the feature map remain minimal. This modified feature map is then processed back into an image. The final image, now containing the hidden code, looks identical to the original when viewed by the human eye. That’s because the code was embedded at the feature map level (not directly into the image pixels), and then turned back into a new image. That’s what makes the watermark invisible to humans (Zhao, 2024).
Think of it like this: imagine an artist painting a picture. Every time they finish a painting, they secretly leave a mark, for example, always making the yellow parts slightly thicker than the rest. No one else would notice this as a hidden signature, but the artist knows it’s there. This subtle thickening of the yellow paint, added after the painting is finished, is like embedding a hidden code into the feature map. Only the artist knows it’s there, and only someone who knows what to look for can recognize it.
The Decoder: An AI That Can “Read” the Hidden Code
Now, in order for the hidden code to be read back, there’s a second AI model involved, called the decoder. Its job: to analyze the image and guess the secret code embedded inside. At first, the decoder’s guesses are totally random. But after training on thousands, or even millions, of samples, it gradually gets better and learns to recognize the hidden patterns.
To make the decoder smarter, during the training process, the images are deliberately altered, cropped, resized, or augmented in different ways. This ensures that even if the image has been slightly edited, the decoder can still read the hidden code (as long as the modifications aren’t too extreme) (DeepMind, 2024).
Let’s return to the earlier analogy. Imagine there’s another person who plays the role of the decoder. The original artist (who always paints with slightly thicker yellow strokes as a secret signature) teaches this person how to recognize the code. The person studies thousands of the artist’s paintings and is asked to guess where the hidden code is. The artist only responds with “yes” or “no.”
At first, the decoder has no idea where the secret code is. But after reviewing so many paintings, the person begins to notice that the yellow parts are consistently thicker. Over time, they figure out that this is the hidden signature. And just like that, the decoder learns how to distinguish between authentic artworks and imitations. In the end, only the artist and the decoder trained to detect that hidden signature can recognize it.
Limitations and Challenges
Of course, the system isn’t perfect. The decoder can only guess the watermark code if the AI model’s creator specifically designed and trained the decoder for that purpose (Kohli et al., 2024). Without it, the video output can’t be detected as AI-generated or real, especially if it’s already highly realistic. That’s exactly why models like Veo 3 have been anticipated and equipped with this method from the start.
Another weakness is that if the image is heavily altered or deliberately injected with noise to fool the decoder, the digital watermark might fail to be detected (Zhao et al., 2024). However, as technology continues to evolve, the hope is that watermarking and detection systems will become more resilient to such manipulations.
What’s also hard to deny is that generative AI is progressing extremely fast, sometimes faster than the detection systems trying to keep up (DeepMind, 2024). Technologies like SynthID must be continuously updated so they don’t fall behind the latest “tricks.” Future AI models might even learn how to bypass or remove watermarks entirely, without leaving a trace.
The Social, Political, and Governance Dimension of AI
From an academic and public policy perspective, this challenge highlights the importance of a multidisciplinary approach and adaptive governance (Daly, 2019). It is not enough to rely solely on technical solutions; what’s also needed are responsive regulations, active civil society engagement, digital literacy education, and independent audit mechanisms to ensure transparency and accountability (Zuiderwijk, 2021). Multi-stakeholder collaboration, among governments, the private sector, academia, and civil society, is key to building an AI governance framework that is inclusive, fair, and sustainable (UN, 2023).
Recent studies have also emphasized that the expansion of AI without strong governance frameworks can undermine public trust and erode the quality of democracy, especially as AI can amplify disinformation, weaken accountability, and intensify social polarization (Kreps & Kriner, 2023). In fact, an empirical study by Chehoudi (2015), which analyzed the relationship between AI development and democratic quality across countries, found that AI adoption without proper oversight tends to negatively correlate with democratic health.
Therefore, the development and oversight of watermarking technologies like SynthID, and other similar platforms, must be embedded within a broader AI governance ecosystem. One that not only champions innovation, but also prioritizes human rights, ethics, and the collective interests of society.
References
BBC News (2024) ‘India election: AI deepfakes and the politics of synthetic reality’. Available at: https://www.bbc.com/news/world-asia-india-68918330 (Accessed: 11 June 2025).
Chehoudi, R. (2025) ‘Artificial intelligence and democracy: pathway to progress or decline?’, Journal of Information Technology & Politics, [online] Available at: https://www.tandfonline.com/doi/full/10.1080/19331681.2025.2473994 (Accessed: 11 June 2025).
Daly, A., Hagendorff, T., Li, H., Mann, M., Marda, V., Wagner, B., Wang, W. and Witteborn, S. (2019) ‘Artificial Intelligence, Governance and Ethics: Global Perspectives’, SSRN Electronic Journal, [online] Available at: https://ssrn.com/abstract=3414805 (Accessed: 11 June 2025).
DeepMind (2024) ‘Watermarking AI‑generated text and video with SynthID’. DeepMind Blog.
Google (2023) ‘AI Principles’. Available at: https://ai.google/principles/ (Accessed: 11 June 2025).
Google (2024) ‘Veo: Google’s most advanced generative video model’. YouTube video], 14 May. Available at: https://www.youtube.com/watch?v=XaDSA8GOy4g (Accessed: 11 June 2025).
Kohli, P., et al. (2024) ‘Scalable watermarking for identifying large language model outputs’, Nature, 615, pp. 123–130.
Kreps, S. and Kriner, D. (2023) ‘How AI Threatens Democracy’, Journal of Democracy, 34(4), pp. 122–131. Available at: https://www.journalofdemocracy.org/articles/how-ai-threatens-democracy/ (Accessed: 11 June 2025).
United Nations (2023) ‘Artificial Intelligence | Office for Digital and Emerging Technologies’, [online] Available at: https://www.un.org/digital-emerging-technologies/content/artificial-intelligence (Accessed: 11 June 2025).
YouTube Official Blog (2024) ‘Our approach to responsible AI innovation’. Available at: https://blog.youtube/inside-youtube/our-approach-to-responsible-ai-innovation/ (Accessed: 11 June 2025).
Zhao, X., et al. (2024) ‘Invisible Adversarial Watermarking: A Novel Security Mechanism for Enhancing Copyright Protection’, NeurIPS.