![]() |
| click to enlarge |
Showing posts with label other. Show all posts
Showing posts with label other. Show all posts
Thursday, April 28, 2016
Entry in #BSidesLV Logo Contest
Here's my entry in the BSides Las Vegas logo contest. The crowd-chosen slogan is "Popping calc.exe since 2009".
Wednesday, March 30, 2016
#Tay Twist: @Tayandyou Twitter Account Was Hijacked ...By Bungling Microsoft Test Engineers (Mar. 30)
This summary is not available. Please
click here to view the post.
Tuesday, March 29, 2016
Media Coverage of #TayFail Was "All Foam, No Beer"
One of the most surprising things I've discovered in the course of investigating and reporting on Microsoft's Tay chatbot is how the rest of the media (traditional and online) have covered it, and how the digital media works in general.
None of the articles in major media included any investigation or research. None. Let that sink in.
All foam, no beer.
None of the articles in major media included any investigation or research. None. Let that sink in.
All foam, no beer.
Sunday, March 27, 2016
Microsoft's Tay Has No AI
(This is the third of three posts about Tay. Previous posts: "Poor Software QA..." and "...Smoking Gun...")
While nearly all the press about Microsoft's Twitter chatbot Tay (@Tayandyou) is about artificial intelligence (AI) and how AI can be poisoned by trolling users, there is a more disturbing possibility:
I say "probably" because the evidence is strong but not conclusive and the Microsoft Research team has not publicly revealed their architecture or methods. But I'm willing to bet on it.
Evidence comes from three places. First is from observing a small non-random sample of Tay tweet and direct message sessions (posted by various users). Second is circumstantial, from composition of the team behind Tay. Third piece of evidence is from a person who claims to have worked at Microsoft Research on Tay until June 2015. He/she made two comments to my first post, but unfortunately deleted the second comment which had lots of details.
While nearly all the press about Microsoft's Twitter chatbot Tay (@Tayandyou) is about artificial intelligence (AI) and how AI can be poisoned by trolling users, there is a more disturbing possibility:
- There is no AI (worthy of the name) in Tay. (probably)
I say "probably" because the evidence is strong but not conclusive and the Microsoft Research team has not publicly revealed their architecture or methods. But I'm willing to bet on it.
Evidence comes from three places. First is from observing a small non-random sample of Tay tweet and direct message sessions (posted by various users). Second is circumstantial, from composition of the team behind Tay. Third piece of evidence is from a person who claims to have worked at Microsoft Research on Tay until June 2015. He/she made two comments to my first post, but unfortunately deleted the second comment which had lots of details.
Saturday, March 26, 2016
Microsoft #TAYFAIL Smoking Gun: ALICE Open Source AI Library and AIML
[Update 3/27/16: see also the next post: Microsoft's Tay has no AI"]
As follow up to my previous post on Microsoft's Tay Twitter chatbot (@Tayandyou), I found evidence of where the "repeat after me" hidden feature came from. Credit goes to SSHX for this lead in his comment:
AIML is acronym for "Artificial Intelligence Markup Language", which "is an XML-compliant language that's easy to learn, and makes it possible for you to begin customizing an Alicebot or creating one from scratch within minutes." ALICE is acronym for "Artificial Linguistic Internet Computer Entity". ALICE is free natural language artificial intelligence chat robot.
As it happens, there is an interactive web page with Base ALICE here. (Try it out yourself.) Here is what happened when I entered "repeat after me" and also "repeat this...":
In Base ALICE, the template response to "repeat after me" is "...". In other words, NOP ("no operation"). This is different from the AIML statement, above, which is ".....Seriously....Lets have a conversation and not play word games.....". Looks like someone just deleted the text following three periods.
But the template response to "repeat this X" is "X" (in quotes), which is consistent with the AIML statement, above.
My assertion about root cause stands: poor QA process on the ALICE rule set allowed the "repeat after me" feature to stay in, when it should have been removed or modified significantly.
Another inference is that "repeat after me" is probably not the only "hidden feature" in AIML rules that could have caused misbehavior. It was just the one that the trolls stumbled upon and exploited. Someone with access to Base ALICE rules and also variants could have exploited these other vulnerabilities.
As follow up to my previous post on Microsoft's Tay Twitter chatbot (@Tayandyou), I found evidence of where the "repeat after me" hidden feature came from. Credit goes to SSHX for this lead in his comment:
"This was a feature of AIML bots as well, that were popular in 'chatrooms' way back in the late 90's. You could ask questions with AIML tags and the bots would automatically start spewing source into the room and flooding it. Proud to say I did get banned from a lot of places."A quick web search revealed great evidence. First, some context.
AIML is acronym for "Artificial Intelligence Markup Language", which "is an XML-compliant language that's easy to learn, and makes it possible for you to begin customizing an Alicebot or creating one from scratch within minutes." ALICE is acronym for "Artificial Linguistic Internet Computer Entity". ALICE is free natural language artificial intelligence chat robot.
Evidence
This Github page has a set of AIML statements staring with "R". (This is a fork of "9/26/2001 ALICE", so there are probably some differences between Base ALICE today.) Here are two statements matching "REPEAT AFTER ME" and "REPEAT THIS".![]() |
| Snippet of AIML statements with "REPEAT AFTER ME" AND "REPEAT THIS" (click to enlarge) |
In Base ALICE, the template response to "repeat after me" is "...". In other words, NOP ("no operation"). This is different from the AIML statement, above, which is ".....Seriously....Lets have a conversation and not play word games.....". Looks like someone just deleted the text following three periods.
But the template response to "repeat this X" is "X" (in quotes), which is consistent with the AIML statement, above.
Conclusion
From this evidence, I infer that Microsoft's Tay chatbot is using the open-sourced ALICE library (or similar AIML library) to implement rule-based behavior. Though they did implement some rules to thwart trolls (e.g. gamergate), they left in other rules from previous versions of ALICE (either Base ALICE or some forked versions).My assertion about root cause stands: poor QA process on the ALICE rule set allowed the "repeat after me" feature to stay in, when it should have been removed or modified significantly.
Another inference is that "repeat after me" is probably not the only "hidden feature" in AIML rules that could have caused misbehavior. It was just the one that the trolls stumbled upon and exploited. Someone with access to Base ALICE rules and also variants could have exploited these other vulnerabilities.
Friday, March 25, 2016
Poor Software QA Is Root Cause of TAY-FAIL (Microsoft's AI Twitter Bot)
[Update 3/26/16 3:40pm: Found the smoking gun. Read this new post. Also the recent post: "Microsoft's Tay has no AI"]
This happened:
I claim: the explanations that blame AI are wrong, at least in the specific case of tay.ai.
This happened:
"On Wednesday morning, the company unveiled Tay [@Tayandyou], a chat bot meant to mimic the verbal tics of a 19-year-old American girl, provided to the world at large via the messaging platforms Twitter, Kik and GroupMe. According to Microsoft, the aim was to 'conduct research on conversational understanding.' Company researchers programmed the bot to respond to messages in an 'entertaining' way, impersonating the audience it was created to target: 18- to 24-year-olds in the US. 'Microsoft’s AI fam from the internet that’s got zero chill,' Tay’s tagline read." (Wired)Then it all went wrong, and Microsoft quickly pulled the plug:
"Hours into the chat bot’s launch, Tay was echoing Donald Trump’s stance on immigration, saying Hitler was right, and agreeing that 9/11 was probably an inside job. By the evening, Tay went offline, saying she was taking a break 'to absorb it all.' " (Wired)Why did it go "terribly wrong"? Here are two articles that assert the problem is in the AI:
- "It’s Your Fault Microsoft’s Teen AI Turned Into Such a Jerk" - Wired tl;dr: "this is just how this kind of AI works"
- "Why Microsoft's 'Tay' AI bot went wrong...AI experts explain why it went terribly wrong" - TechRepublic tl;dr: "The system is designed to learn from its users, so it will become a reflection of their behavior".
The "blame AI" argument is: if you troll an AI bot hard enough and long enough, it will learn to be racist and vulgar. ([Update] For an example, see this section, at the end of this post)
Wednesday, April 15, 2015
Entry in Schneier's Eighth Movie-Plot Threat Contest
Every year, on April Fool's Day, Bruce Schneier hosts a "movie plot threat" contest on his blog. This year's theme is "evils of encryption". This is my third year submitting an entry (I won two years ago -- w00t!). Here is my entry for the 8th contest (500 word limit):
Thursday, May 1, 2014
Splitting this blog and moving to Octopress
I've decided to split this blog to separate my academic posts from my industry posts. I'm going to be blogging more about my dissertation and related works in progress, and I suspect that most of my industry readers won't be interested and I don't want to dilute my posts on industry topics -- information security, risk, performance metrics, etc.
Google's Blogger has worked well for me, but I've decided to move to Octopress. I'll spare you all the details of the decision process but here's a post that describes the process and benefits. I'm also following in the footsteps of others in my community (e.g. Securitymetrics.org, Data Driven Security, and Adam Elkus).
The industry blog with be renamed "Meritology Blog" and will have a meritology.com URL. The academic blog will be "Exploring Possibility Space" and will have an exploringpossibilityspace.com URL. I aim to move all the Blogger posts to these so that the archives are available under both.
I'll let you know when this goes live, and hopefully there will be redirection once the move is complete.
Google's Blogger has worked well for me, but I've decided to move to Octopress. I'll spare you all the details of the decision process but here's a post that describes the process and benefits. I'm also following in the footsteps of others in my community (e.g. Securitymetrics.org, Data Driven Security, and Adam Elkus).
The industry blog with be renamed "Meritology Blog" and will have a meritology.com URL. The academic blog will be "Exploring Possibility Space" and will have an exploringpossibilityspace.com URL. I aim to move all the Blogger posts to these so that the archives are available under both.
I'll let you know when this goes live, and hopefully there will be redirection once the move is complete.
Thursday, April 17, 2014
"Creative Destruction": 500 word entry for Schneier's Movie Plot Contest
Since I won last year, I wasn't going to enter this year. But my imagination started turning and this came out. Hope you enjoy it.
My entry:
Bruce Schneier's 7th Annual Movie Plot Contest
Theme: NSA wins! But how? (full description and all entries are here)My entry:
Creative Destruction
June 2014 – March 2015: Stock market booms.
June 2014: Snowden revelations trigger international political scandals.
July: Feinstein-Rogers Intelligence Reform Bill passes, breaks up NSA. “Largest garage sale in history”.
Headline: “NSA Nuked”
August: 10,000 NSA workers are laid off.
September – December: Open Source projects, Working Groups see influx of volunteers.
September – November: “NSA garage sale” draws small contractors and public-private partnerships spread over 50 states. Private equity firms are buyers – Flatiron Partners, Narsil Capital, and Tech Disruptions.
September – November 2014: Flurry of privacy and security scandals hit big firms. Lawsuits, investigations, and criminal indictments follow.
November: “Alt Apps Group” formed: “Secure, private, and ad-free”. Most members are majority owned by Flatiron, Narsil Capital, or Tech Disruptions.
July – December: Symantec goes on buying spree: Webroot, Cloudflare, StackExchange, Disqus, Rapid7, and MaaS360 – all funded by Flatiron, Narsil Capital, or Tech Disruptions.
December: Puerto Rico Bridge Initiative announces completion of 50GB fiber optic cable.
January 2015: Private equity firms, led by Flatiron, make offer for Symantec.
February: Loren Reynolds, rookie Equity Analyst at JPMorgan Chase working on Symantec project, is accidentally copied on email from Flatiron:
“Confirming that Launch has been accelerated to March. Don’t use email anymore.”Loren is puzzled by the distribution list:
March: Traffic and membership surges at Alt Apps Group members. Fatherly achieves 70% share in Certificate Authority market.
- Puerto Rico Bridge Initiative (PRBI)
- Economic Development Corporation Utah (EDCU)
- Fatherly (formerly GoDaddy)
- DuckDuckGo
- Safebook (startup)
March: Loren receives email from friend Zoltin, networking expert:
“PRBI isn’t 50GB. It’s 75TB – the highest capacity in the world!!! WTF! There’s more. Same capacity cables to Bermuda and to Azores. Looks like they are bypassing US-Europe cables. Big news: Big data center on PR now complete.”April: Google earnings due 4/15; then Facebook, Apple, Twitter, Microsoft, and Verizon on 4/22.
April: Over beers, Loren hears rumors of large short positions in technology stocks from a few “weird” hedge funds.
April 4: Loren receives email from Sarah:
“EDCU is gatekeeper on NSA Utah Data Center.”
April 15: Google disappoints. Earnings down 40% on flat revenue. Stock falls 30%, overall market falls 10%.
April 20: Loren discovers link between former NSA executives and Flatiron, Narsil Capital, and Tech Disruptions. Finds NSA people behind tech firm scandals in the previous Fall 2014. Tip of the iceberg, she suspects.[This has a few lines added, so it's beyond 500 word limit. But the entry on Bruce's site is below the limit.]
April 21: Loren sends IM to her husband, an Assistant DA in the Southern District of NY:
“Must see you ASAP. NSA & GCHQ live on. They’ve gone legit – running private businesses and funds. THEY ARE KILLING THE INTERNET AD BUSINESS.”After clicking “send”, her computer freezes. She reaches for her smart phone to call her husband, but the directory is empty. She dials the number manually, but gets an “out of service” signal, followed by “low battery”. The phone dies.
Running down seven flights of stairs, Loren races to her car. She jumps in, starts the engine, and backs out with a screech. One turn from the exit the car engine suddenly cuts out and brakes lock. The car crashes into a cement pillar. The airbag fails to deploy. Loren is out cold.
Monday, March 10, 2014
Boomer weasel words: 'high net worth individuals' as euphemism for 'rich people'
![]() |
| (Source) |
(You might compare this to my previous post regarding "Baby on Board" signs.)
The following graphs are from Google Ngram Viewer, which show relative frequency of word phrases in American English books up to the year 2008. Notice that the phrase "high net worth individuals" appears first around 1980.
![]() |
| (click to enlarge) |
Monday, February 24, 2014
#BSidesSF Prezo: Getting a Grip on Unexpected Consequences
Here are the slides I'm presenting today at B-Sides San Francisco (4pm). I suggest that you download it as PPTX because it is best viewed in PowerPoint so you can read the stories in the speaker notes.
Monday, January 13, 2014
Guest on "Data-driven Security Podcast" Ep. 1
I was a guest on the new Data-driven Security Podcast, episode 1. There's the usual audio and also a video (1 hour 15 minutes). Along with hosts Bob Rudis and Jay Jacobs, I joined Michael Roytman and Alex Pinto for a lively conversation about how we all got into the data analysis side of information security and where we see it going.
The podcast and also the web site and blog are associated with a new book with the same title, Data-driven Security, authored by Bob and Jay. I was technical editor, so I can honestly say that I've read the whole book. I heartily recommend it to any information security professional or manager. It is a perfect "on-ramp" into data science and visualization as applied to information security, and it's written in your language.
The podcast and also the web site and blog are associated with a new book with the same title, Data-driven Security, authored by Bob and Jay. I was technical editor, so I can honestly say that I've read the whole book. I heartily recommend it to any information security professional or manager. It is a perfect "on-ramp" into data science and visualization as applied to information security, and it's written in your language.
Tuesday, October 1, 2013
Pantyhose attitudes correlate to GOP's "Suicide Caucus" districts
The New Yorker blog has an interesting post on the geography and demography behind the current government shutdown: "Where the GOP's Suicide Caucus Lives". Read the article because it has a lot of good data on the districts of 80 Representatives who have pushed the Republican leadership into this battle. (The label "Suicide Caucus" was coined by conservative commentator Charles Krauthammer.)
Here's the map of districts of those Representatives (click to enlarge):
This immediately reminded me of the map I posted regarding attitudes toward pantyhose in this post. Here's the map from that post:
Just to see if the two might be correlated, I physically combined the maps (click to enlarge):
[Edit 10/8/13 -- added this map, below]
Here's the same data, but this map only shows the relative prevalence of people who believe it is "acceptable". Red indicates very high prevalence, yellow indicates medium prevalence, and blue, low prevalence. (Click on map to enlarge.)
I think the correlation is striking, though certainly not perfect especially in certain regions. For example, the region of East Texas and West Louisiana is a stronghold of the Suicide Caucus but is very pro-pantyhose. Likewise, we'd expect Suicide Caucus members in Nebraska, Oklahoma, South Dakota, and West Virginia, but there are none (at least indicated by signatures on the letter). Also, the most pro-pantyhose region of Colorado is in the Caucus, and likewise with Idaho.
Food for thought and grounds for further examination.
Here's the map of districts of those Representatives (click to enlarge):
This immediately reminded me of the map I posted regarding attitudes toward pantyhose in this post. Here's the map from that post:
Just to see if the two might be correlated, I physically combined the maps (click to enlarge):
[Edit 10/8/13 -- added this map, below]
Here's the same data, but this map only shows the relative prevalence of people who believe it is "acceptable". Red indicates very high prevalence, yellow indicates medium prevalence, and blue, low prevalence. (Click on map to enlarge.)
I think the correlation is striking, though certainly not perfect especially in certain regions. For example, the region of East Texas and West Louisiana is a stronghold of the Suicide Caucus but is very pro-pantyhose. Likewise, we'd expect Suicide Caucus members in Nebraska, Oklahoma, South Dakota, and West Virginia, but there are none (at least indicated by signatures on the letter). Also, the most pro-pantyhose region of Colorado is in the Caucus, and likewise with Idaho.
Food for thought and grounds for further examination.
Sunday, September 22, 2013
Blogging as rapid prototyping
People blog for a lot of reasons and in many styles. Except for occasional posts like this one, I rarely write short posts. No doubt, some potential readers will find this to be a big turnoff. I'm fine with that. I know who I aim to serve and who I don't. This blog isn't for people who have short attention spans or who only want bit-sized "nuggets".
In addition to writing to serve readers, I also write for my own purposes. I've discovered that blogging works best when it serves as a way to rapidly prototype ideas and methods that might later become academic papers, industry presentations, book chapters, books, software models, and such. These final products take a lot of time and effort to produce in final form. Blogging gives me the opportunity to get started on them, focusing on just a few ideas at a time and without the need to have everything worked out. Plus I get to see how people react, either through page views, social media comments, blog comments, or private email.
In addition to writing to serve readers, I also write for my own purposes. I've discovered that blogging works best when it serves as a way to rapidly prototype ideas and methods that might later become academic papers, industry presentations, book chapters, books, software models, and such. These final products take a lot of time and effort to produce in final form. Blogging gives me the opportunity to get started on them, focusing on just a few ideas at a time and without the need to have everything worked out. Plus I get to see how people react, either through page views, social media comments, blog comments, or private email.
Thursday, September 12, 2013
I'm leaving Facebook (Frog escapes slowly boiling pot)
![]() |
| That's a frog on the handle. It was in the pot but jumped out when things got too hot. |
I'm leaving Facebook this week -- permanently. I'm tired of the creeping encroachments on my privacy. Also I'm no longer willing to be a part of Facebook's quest to commercialize and make public all of our social relations and interactions.
The most recent privacy policy changes are the proximate cause (see this, this, this and this). Though protest and government scrutiny have prompted Facebook to delay implementation, the trend is clear.
The title of this post refers to the story of the Boiling Frog:
If you drop a frog in a pot of boiling water, it will of course frantically try to clamber out. But if you place it gently in a pot of tepid water and turn the heat on low, it will float there quite placidly. As the water gradually heats up, the frog will sink into a tranquil stupor, exactly like one of us in a hot bath, and before long, with a smile on its face, it will unresistingly allow itself to be boiled to death.
(version of the story from Daniel Quinn's The Story of B)I'm not against businesses making money through advertising in their "free" services. It's just the way Facebook is doing it that deeply bothers me.
Privacy isn't just "not disclosing private information". It's also about people keeping control of their private information, where and how it is used, and by whom. Facebook's latest changes are forcing users like me to give away vital elements of control, in my opinion.
Finally, I don't trust them to keep to the spirit of privacy. Facebook's definition of privacy is like Bill Clinton's definition of "sexual relations" -- an unreasonably narrow definition whose rhetorical aim is to dissemble. At best, I believe Facebook will continue to keep to the letter of their constantly shifting privacy policy and user agreement, all the while constantly finding ways to subtly erode our privacy. At worst -- well, obviously very bad things would happen. But I'm acting on the assumption of the best case, not the worst.
Bye, bye, Facebook. And I won't be coming back.
Saturday, July 20, 2013
Let’s Shake Up the Social Sciences
Given that I'm a student in the first-ever Department of Computational Social Science, I strongly agree with Nicholas Christakis in his New York Times article "Let’s Shake Up the Social Sciences". I especially like the points he makes regarding teaching social science students to do experiments early in their education process.
Friday, July 19, 2013
Guest on the Risk Science Podcast
On Episode 3 of the Risk Science Podcast, I had a nice conversation with my friends Jay Jacobs and Ally Miller. The topics included on the balance in simplifying complexity, the need to get more industry people involved in the WEIS conference (as participants and presenters), writing winning movie plots about cyber war, and the learning curve for R.
In case you don't recognize it, the Risk Science Podcast (@Risksci on twitter) is new and improved. Previous iterations were the Risk Hose Podcast, and before that, the SIRA Podcast.
In case you don't recognize it, the Risk Science Podcast (@Risksci on twitter) is new and improved. Previous iterations were the Risk Hose Podcast, and before that, the SIRA Podcast.
Saturday, June 22, 2013
Let the explorations begin
I think it's time to create a personal blog. For a few years, I've been blogging at New School of Information Security about information security, especially risk estimation, metrics, and risk management. I've used Twitter (@MrMeritology) for these and other topics, including various topics related to my PhD program and dissertation. I'll still blog at NewSchool from time to time, and I will continue to use Twitter for community interaction and resource sharing.
The blog format allows for longer essays and maybe novel content (e.g. interactive animations). I imagine this blog will be more open in terms of topics and ideas that interest me as I come across them. I'll also be blogging my progress through my dissertation and my experiences with various tools and methods as I learn them.
The content will be mostly professional, intellectual, and technical, with some occasional philosophy and personal essays.
I won't be blogging personal or family stuff.
The blog format allows for longer essays and maybe novel content (e.g. interactive animations). I imagine this blog will be more open in terms of topics and ideas that interest me as I come across them. I'll also be blogging my progress through my dissertation and my experiences with various tools and methods as I learn them.
The content will be mostly professional, intellectual, and technical, with some occasional philosophy and personal essays.
I won't be blogging personal or family stuff.
Subscribe to:
Posts (Atom)












