The machine learning WAF

Last week I was at Navaja Negra Conference 2023. Ruben Garrote García and Rubén Ródenas Cebrián gave a talk on how to bypass a WAF with payloads generated by machine learning.

I liked the approach. Thanks, guys.

Part of what they used was a project that Sergio Daniel Fernández Atochero and I built when we worked together at BBVA Innovation Labs: WAF-Brain.

The WAF-Brain project

In WAF-Brain we trained a model with a lot of real SQL injection attacks, the ones SQLMap and OWASP ZAP actually throw at you.

The question was whether you could take real attacks from real tools, train a model with them and get a useful detector out of it. You can. We did.

Then we put our PoC next to ModSecurity, SQL injection only, and a few things came out of that:

  • ModSecurity detected worse than WAF-Brain in most of the attacks.
  • WAF-Brain was faster. And its detection time was almost constant, which ModSecurity’s was not.
  • Our model still had work to do with very short and very long payloads.

WAF-Brain was meant to be a starting point for other people’s research, not a finished product.

If you want to know more, the repo has a research section.

The butterfly effect

What surprised me was hearing them say that WAF-Brain had been the base of several research papers. A few of them:

We tend to forget that what we build can end up inspiring someone on the other side of the world.

Sometimes the butterfly effect is a nice thing.