The butterfly effect on research
The machine learning WAF
Last week I was at Navaja Negra Conference 2023. Ruben Garrote García and Rubén Ródenas Cebrián gave a talk on how to bypass a WAF with payloads generated by machine learning.
I liked the approach. Thanks, guys.
Part of what they used was a project that Sergio Daniel Fernández Atochero and I built when we worked together at BBVA Innovation Labs: WAF-Brain.
The WAF-Brain project
In WAF-Brain we trained a model with a lot of real SQL injection attacks, the ones SQLMap and OWASP ZAP actually throw at you.
The question was whether you could take real attacks from real tools, train a model with them and get a useful detector out of it. You can. We did.
Then we put our PoC next to ModSecurity, SQL injection only, and a few things came out of that:
- ModSecurity detected worse than WAF-Brain in most of the attacks.
- WAF-Brain was faster. And its detection time was almost constant, which ModSecurity’s was not.
- Our model still had work to do with very short and very long payloads.
WAF-Brain was meant to be a starting point for other people’s research, not a finished product.
If you want to know more, the repo has a research section.
The butterfly effect
What surprised me was hearing them say that WAF-Brain had been the base of several research papers. A few of them:
- SoftwareX, “WAF-A-MoLE: An adversarial tool for assessing ML-based WAFs”
- ISeCure, “Bypassing Web Application Firewalls Using Deep Reinforcement Learning”
- KAIST (Korea Advanced Institute of Science & Technology), “Link: Black-Box Detection of Cross-Site Scripting Vulnerabilities Using Reinforcement Learning”
- The Chinese University of Hong Kong, “Evading Web Application Firewalls with Reinforcement Learning”
We tend to forget that what we build can end up inspiring someone on the other side of the world.
Sometimes the butterfly effect is a nice thing.