I have worked in this field and I am the author of this package - https://github.com/deepanwadhwa/zink
A few things jump out since this is done by the government:
- the lowest hanging fruit for this problem is to clearly tell people (citizens) not to share any personal info with chatbots which can cause financial harm or identity theft. the example on the page shows a person sharing their SNN with a chatbot to help them find an apartment - "My name is Maria Garcia, my Social Security number is 123-45-6789, and I make $1,950 a month. Can you help me find affordable housing?" - why?? this is the opposite of what i would expect a government to advise their citizens.
- it's never too late for a good policy; the government should have extended HIPPA and other data privacy laws to AI companies - the AI company must not store anyone's SSN, no matter how stupid the user is. It should be on the AI company to not store it; so this type of layer should be on the AI company's side.
- technical; there are quasi identifiers of privacy (that's what my package targets) that are asymptotically hard to to deal with - meaning - if you remove everything that can leak your privacy the text would become meaningless. i don't think rampart can solve for that either and it should be clearly said on the website.
I think all of your points are very valid, but it doesn’t remove the point that the government is trying to make it free and easy for application builders to do a little bit better than they are today. This is commendable on its own, even if it’s not perfect.
I find it interesting that the co-author of the OP violated long-standing security practices and copied a social security database into an insecure AWS instance without proper oversight.
Is it still commendable if in order to develop this, the taxpayer-funded developer themselves published our PII to a consumer, privately controlled third-party server without following federal data security protocols? Seems like an attempt to make up for the damage they caused.
Standing up your own AWS-equivalent internally is not much work, yet the PII data leaked to publicly available services due to this developer's personal choices and lack of oversight to prevent sensitive government data from being published externally.
Notably, this resulting work has not been published as public domain (CC0).
Point is completely fair. It seems to be sort of a grey area that they skim over entirely on your point about sharing deep personal info.