Claim 7 in this patent application describes how depth sensors are used as part of an image authentication process, which would make such a workaround more difficult:

https://image-ppubs.uspto.gov/dirsearch-public/print/downloa...

The Apple Reference Image feature is here launched on iPhone 18 Pro and iPhone 18 Pro Max that both have built-in LiDAR sensors that could be used for this process.

Apple's current implementation doesn't integrate LiDAR. And LiDAR wouldn't be enough here, it's trivial to block the projector and hide the dot pattern. No dot pattern = iPhone thinks the object is far away, which is what happens in landscape photos.

A better fix is to take photos with all three iPhone cameras simultaneously, ideally as a 2-3s video, and use the parallax/multiple perspectives to extract depth information. The video files (Possibly audio too) could also be included with the verified image as additional verification.

They can also prevent photos if iPhone detects the LiDAR sensor is covered, similar to how Meta does it with their camera glasses.

> it's trivial to block the projector and hide the dot pattern. No dot pattern = iPhone thinks the object is far away, which is what happens in landscape photos.

I've never looked at the LiDAR hardware, but where is the emitter in relation to the receiver. Why would the LiDAR not reflect off of whatever you're blocking it with and return a very short flight meaning it was very close?

I think the plan is to block the emitter, not the emitter and the receiver together

Paint the stopper vantablack then.

I knew keeping a pot of vantablack in my jacket pocket would come in handy.

> all three iPhone cameras simultaneously, ideally as a 2-3s video, and use the parallax/multiple perspectives to extract depth information.

I think optics could be used to make each camera see a different image.

A video could show shake, which could be verified against readings from the phone's accelerometer -- but you could just hold it still and claim that it was on a tripod.

I don't think it's an either or - additional data signals that need to correlate to authenticate will increase confidence. You can use multiple other signals to evaluate whether something is truly a landscape photo, and in that case not require a LiDAR capture, but if you are inside and at close range then you could assume that it should be part of scoring the authentication.

Similarly, LiDAR alone will help disqualify cases where someone is just taking a picture of e.g. a landscape target of the Golden Gate, but that it shown on a screen 1 meter away.

Right but I’d argue that realistically this feature is going to be most useful when taking photos of things reasonably close by, people especially, rather than landscape photography.

This approach makes me wonder if the future is actually going to move towards visual cryptography.

Um… that's already a thing [1]:

    TL;DR: What is C2PA in 60 seconds

    What: An open technical standard for embedding cryptographically signed provenance data inside digital media files.

    Who: Created by a coalition founded by Adobe, Arm, BBC, Intel, Microsoft, and Truepic in February 2021.

    How: A C2PA Manifest (also called a Content Credential) travels inside the file and records who made it, when, and what tools were used.

    Why: Deepfake incidents surged from 500,000 to 8 million cases between 2023 and 2025. Provenance gives media a verifiable chain of custody.
[1]: https://c2paviewer.com/articles/what-is-c2pa

iPhone lidar only works up to like 16 feet in the easiest lighting conditions (indoors) and may be functionally ineffective outdoors.

Still, that means that either the fake target scene and your screen presenting it would need to be outside of LiDAR sensor bounds, or you'd need to find a way to make the depth sensor data conform with your fake scene, both increasing the difficulty of producing a forgery.

https://news.ycombinator.com/item?id=49721878

> increasing the difficulty of producing a forgery

The problem with this thinking is twofold:

1) Whether it actually meaningfully increases the difficulty of a forgery remains to be seen. Despite their initial language about discerning real events, we see no details here about what scene information is used.

2) It increases the potential value of a forgery because now your forgery is attested by Apple.

So it either makes it easier to defraud people or more worthwhile to put in the effort to defraud people or both. None of those outcomes are great.

So you need a big enough screen to cover the entire field of view at 16 ft? Sounds expensive

Or, probably just some optics, like a slanted $9 IR mirror [1] in front of it, to direct the lidar to the sky/absorption box. Then you can point at the high res HDR TV that's probably already in your living room.

[1] https://commonlands.com/products/ir-cut-filters-csp650?srslt...

Well, first, that's only expensive if you're poor. The world is absolutely full of people who can easily piss away your entire annual income throwing a house party.

But I really mean that if the lidar barely works outdoors anyway then actually you don't need to be 16 feet away at all.

Anyway, one may presume that they've thought about this.

Thought about it and also are bright enough not to fall for the “if a single person dies wearing a seat belt, we should abandon seat belts because they do no good at all” fallacy.

It’s almost certainly possible to fool v1 of this system, for some images, in some contexts. It would be shocking if the first implementation was completely perfect. But maybe it’s better than nothing?

The problem is that if defeating it is trivial, then it _authenticates_ fake images.

The problem is that it makes it easier to fool people and provide "cryptographic" evidence of validity, backed by big tech.

It's purpose is to stop bad actors from passing of fake as real just as much as it is to prevent real images being dismissed as fake.

It's bad to begin with that Apple should be the arbiters of reality.

> It’s almost certainly possible to fool v1 of this system, for some images, in some contexts. It would be shocking if the first implementation was completely perfect.

Knowing Apple, they've been working on and testing Apple Reference Image for years.

It being perfect isn't the issue; it's that random people on the internet who are just learning about this assume Apple's engineers haven't already thought about everything (and more) mentioned in this thread.

> It being perfect isn't the issue; it's that random people on the internet who are just learning about this assume Apple's engineers haven't already thought about everything (and more) mentioned in this thread.

Given how many bugs there are in macOS and how long they have remained there, I (who have been writing iOS apps from the release of the first retina iPod touch until AI got good) functionally agree with such people; at best, I think Apple's engineers haven't actually solved everything (and more) mentioned in this thread, even if every one of these things may have come up in discussions and even reached an official backlog or task list or similar.

If they thought about it, then why isn't this very obvious failure point even mentioned in the technical breakdown?

> But maybe it’s better than nothing?

I think this will depend on how it gets used. I can imagine numerous outcomes where it's in fact worse than nothing (significantly more effective blackmail, for instance).

> maybe it’s better than nothing?

While that is not quite my bar of confidence when implementing wide-reaching technologies that have numerous unexplored knock-on effects, I guess the calculus must have been different on Infinite Loop recently.

I like that seatbelt argument example.

I’ve seen that type of argument a million times, and I’ll certainly reuse that.

Put some sugar on it. Please.

Def Leppard, is that you?

furthermore, couldnt you do parallax from the multiple cameras as well as flicker the flash?

seems pretty easy to make it sufficiently difficult to trick the system