My First Capture The Flag

I played my first Capture the Flag when I was in Kosovo. Not the game you play in the woods (always a favorite of mine growing up) or in Counter Strike, but as a hacking contest. The organizers build a set of deliberately broken programs, hide a secret string inside each one, and you score points by breaking in and reading it (the string is the flag). I didn’t mean to play, I was just curious how it worked…

Of course I have no idea how to code, much less do any hacking, but that’s where my AI abilities come in… because as you’ve no doubt seen in the news recently it’s very good at finding vulnerabilities. And it turns out it’s also very fun.

Recruitment

This was a FLOSSK (Free Libre Open Source Software Kosova) event as part of their annual conference I was speaking at in Prishtina.

I missed the kick off of the event and first realized it was happening when I ran into Ruhan, who has a volunteer for the conference, but also running some software on his computer between checking people in.

He was running Claude and explained he was running out of tokens but was competing in the CTF challenge. I remembered meeting some folks over the summer that were travelling and competing in regional CTF events and I asked him how it worked.

Ruhan is one of youngest members of FLOSSK, and I had met him over the summer as well. He showed me how to login and then said I could just join his team, “Legends,” since he was playing solo. I had some time to kill before the next weminar i wanted to attend, and an excess of claude tokens, so I joined his time to try to figure out how it worked.

How to CTF

A Capture the Flag is a hacking competition with training wheels and guardrails. The organizers stand up small programs that are deliberately broken in some specific way, and hide a secret string inside each one, called the flag. Your job is to find the weakness, pull the flag out, and paste it into the scoreboard for points. Everything you attack belongs to the organizers and is meant to be attacked. The target is the little broken program they gave you, never the scoreboard itself or anyone else’s machine. I was told that “creative solutions” were frowned upon.

A particular flag’s point value is not fixed. The more teams crack a given challenge, the less it pays out, for everyone, retroactively. So the figure next to each flag here is what it was worth at the moment I solved it, and by the final whistle it had usually drifted lower. I didn’t realize this until doing an analysis after the game, and this led me to chase higher numbers rather than focus on just completing a flag and moving on to the next one.

I Don’t Know How to CTF

Obvioulsy I was not sitting there writing exploit code. I had an AI coding assistant running and my job was to respond to it when it was dumb with my own dumb ideas. My method was to paste in each challenge, wait for it to fail, and then ask questions to try to understand what was happening. The AI wrote and ran the actual payloads, and told me what the results were. When something failed, I read its report, changed the tactic, and we went again. I did provide my AI coder quite a bit of ideas, strategy, and hunches (and stubbornness, it kept wanting to give up!) which I assume is becaise Anthropic wouldn’t let me use Fable or even Opus for the challenge because it was “hacking” adjacent. In fact, it wouldn’t even let me use the more advanced models to write up a summary of what happened!

If that sounds like cheating, that is because it is. Initially the rules of the game did not mention any ban of AI tools, and we believe most of the teams were using them like us. Hard to know for certain. I had no idea what was going on, and this is how I build everything else, so it’s the only way I was even able to participate.

What was initially frustraying but eventually kind of fun was how much skill is left once you hand off the coding… ideas about what to try next was most of the game, and that kept me engaged for the rest of the conference day.

Chatbot 2000

The first challenge I picked was a chatbot hiding a secret code, and the game was to talk it into leaking the code past a filter. This felt like something I was born to be able to do. But the server behind it kept dying. Every message hung for two minutes and then failed. I relaunched it, it died again. At first I was worried I had done it, that I had somehow crashed the bot for everyone by making too many requests with AI. Maybe I had? I don’t think so now, but still don’t know for sure. I\ spent points on a hint, because I was confused how the whole process worked, so for a good while the only thing I was contributing to my Legends team was a negative score. That felt bad.

We eventually found, with evidence rather than a bad feeling, that the failure was on their end and not ours (caused by me? I dont’ think so…). The front page loaded instantly while the chat endpoint hung, which meant the web app was healthy and only the model behind it was stuck. So we reported it to the organizers and moved on. They ended up swapping the broken challenges out entirely. I was disappointed because I was excited to do some “social engineering” on the chat bot.

The Sandbox

I finally solved one! It was called template-sandbox-escape. The setup: a tool for maintainers to preview release notes using a tiny template language, the kind of thing where you write {{ version }} and it fills in the number. The classic way these get hacked (obv I looked this up after) is to type code into the template instead of text and have the server run it. The authors knew that, so they wrapped it in a sandbox and, in the challenge description dared us to get OUT of the sandbox.

I let my little Claude friend go… First it confirmed the sandbox really was running the input. It sent {{ 7*7 }} and the sandbox answered 49. From there it was about mapping the walls they had built, and there were three…

NOTE: I’m going to say “I” for what Claude and I did “together”. If it’s technical, I’ll try to leave it as just Claude. But if it’s technical, it was Claude.

A usual way out of one of these sandboxes leans on special Python names wrapped in double underscores, like class written as __class__. Wall 1 bans that text outright. The other walls remove the filters and shortcuts you would reach for next, and they empty the room so the only thing I had to work with was a single object called release, whose only useful piece was the version number, a plain string.

The Way Out

The key was a quiet feature of Python text formatting. You can write a template like "{0.version}".format(release) and it reaches into the object and pulls out the field for you. The important part I learned (from Claude): that formatting does its own reaching… through a side door the sandbox was not watching. So I could get at the forbidden names after all, as long as I could spell them without typing the banned double underscore in my text.

That spelling problem had a clever answer. I never typed two underscores. I built them at the moment of running, by taking one underscore and repeating it, then gluing the pieces together. The text that was sent was clean. The text the server assembled a split second later was the forbidden word.

Put together, the first key that proved the escape looked like this. It asks the object for its class, with the double underscores built on the fly:

{{ "".join(["{0.", "_"*2, "class", "_"*2, "}"]).format(release) }}

The server answered <class '__main__.Release'>. That was the moment we were through the wall. From there Claude could walk from the object to the program’s own back office, the place where it keeps all its variables.

The Flag

We pointed the same trick at the program’s internal list of variables and asked it to print the whole thing out. Buried in the dump was a line that told us the challenge had been handing over the answer the entire time, if I could only reach it.

Captured · 550 points
SFK26{8e520efa88b0f37c7246b989bec89451}

I literally ran over to Ruhan. My first win!

The program’s own notes admitted the design out loud. It said the flag was read straight from the environment, and that it had tried to lock the door with a word blocklist and a ban on double underscores.

Second Flag: Bug Tracker

A few tasks later, after the organizers swapped out the challenges that kept breaking, a new one appeared in the pwn category, the one for breaking programs instead of websites. It was a small service called the FLOSSK Bug Tracker. You connect to it, type report and a message, and it files your bug. The description was almost a confession: whoever wrote the report handler trusted the message a little more than they should have.

That single line is the puzzle… The bug tracker did not just print my message, it treated my message as the instructions for how to print. Turns out this is an old and famous mistake called a format string bug. It’s like handing someone a note to read aloud, and they read not only your words but any stage directions you tucked between them.

In this case, plain text does nothing. But a handful of special codes, slipped into the note, make the “someone” literally act: one code makes them read a secret off their own desk, and another, the dangerous one, makes them write a value back onto the desk, wherever you point.

Finding the bug was one message. Claude sent a report stuffed with those read codes, and the program printed back a stream of its own internal numbers, raw memory that a bug report should never be able to see. So we knew the door was open.

The goal was narrow. Somewhere in the program’s memory sits a single yes/no switch named is_admin, and it starts on no. The program has a command, whoami, that checks the switch: if it is no, you are told you are a guest; if it is yes, the program hands you the flag. So the whole challenge came down to reaching in and flipping one switch from no to yes, using nothing but a bug report.

Finding the exact spot to flip was not something I knew how to do. The challenge also handed out a copy of the program itself, and Claude and I ignored that gift for far too long, trying to guess the switch’s location by trial and error (me and Claude LOVED brute force attempts). Each wrong guess leaned on the little server until it fell over, and I had to wait for it to come back. And this whole time we’re also working on a dead line. The right move, once I finally yelled at Claude enough to find it, took about thirty seconds: open the program file, which still had all its parts labeled, and look up exactly where is_admin lives. It sat at address 0x404120. With that, the finishing move was one carefully built report:

report %c%16$hhn + padding + the address 0x404120

It printed one character, counted everything it had printed so far, and used that count as the value to stamp onto the switch. Then Claude typed whoami.

The program checked the switch, found it flipped to yes, and printed the flag.

Captured · 642 points
SFK26{9384fcf707b9b27fef1e7ec04aa52c31}

The lesson from this one was the opposite of being clever. If I had paid attention to what I was handed we could have resolved it much faster. The answer was sitting in the program file the whole time, labeled in plain sight, and my reflex to brute-force past it cost me far more than the real break-in did.

Third Flag: A Door and a Keyhole

The third one I cracked was also in the pwn category. It was a little network program hiding a flag it was never meant to give up. This is the kind of bug that seems impossible until you watch it happen (or in my case, let Claude tell you what happened). We typed far more text than the program expected, and the overflow let us change where the program went next.

Picture the program’s memory as a stack of index cards. One card holds something special: the address the program will jump back to when it finishes what it’s doing. You are only supposed to write inside your own card. But this program copied my input into a card without checking the length, so when I kept typing, my text ran off the edge and wrote over the “where was I?” note on the card below. Change that note, and when the program finishes, it jumps wherever I say.

Where I wanted it to jump was a function the authors had left inside the program called win, which prints the flag but never normally runs. The whole job was to overwrite the return note with the address of win.

Two things the authors left undone made this work. There was no stack canary, a tripwire value the program checks before it trusts that return note, so nothing noticed my overflow. And the moment I connected, the program printed one of its own internal addresses, a leak. Programs normally shuffle where they live in memory each run to make this exact attack hard. From that single leaked number we could work out where win sat this time.

There was one last snag… Claude pointed the return note straight at win and the program crashed. The reason is that processors insist the memory be lined up on an even boundary before certain instructions run, and jumping directly into win landed one step off. The fix was to jump first to a single do-nothing instruction, a lone ret, which nudged everything back into line, and then on to win. Clever.

Captured · 578 points
SFK26{8b1529eb9dc77b15c913b51c0ea5cc51}

Timing Matters (and so do hints)

Not every door opened. The one I failed was a web challenge we almost had, about login tokens, the little signed passes a site hands you so it knows who you are on your next click.

This site used two kinds of pass. Visitor passes were signed with a proper lock-and-key pair, where the site keeps a private key and anyone can hold the public half to check a signature. “Internal” passes were trusted a different way, signed with a shared secret word that only the site was supposed to know. The mistake was quiet and fatal: the site used the public half of the visitor key as the shared secret for the internal passes. The public half is, by design, printed for everyone. So we could write our own internal passes and sign them with a secret that was never secret, and the server accepted them as genuine.

We beat the cryptography, which we thought was the genuinely hard part, and then lost to the clock… The only step left was one exact label the pass had to carry to satisfy a final check, and it sat behind a paid hint we chose not to buy (I wanted to, but Ruhan hated spending points to get hints). Time ran out with us one word from the door. Would we have figured it out in time with the hint? Maybe… maybe not.

Conclusion

Obviously I didn’t learn much about hacking (except that Anthropic is very paranoid about using the advanced models). But I did learn how to direct the LLM at the problem. I understand now, in a way I did not before this, what a sandbox is trying to stop, why a program should never treat your input as instructions, and how one misaligned step can crash an attack that is otherwise perfect.

I did not learn that by studying or even understanding the problem… I learned it by trying to grasp the concept, chatting with Claude, then having an idea, watching the LLM try it, and seeing exactly where it broke. And I was surprised how much fun that is! Reading a challenge, guessing at its soft spot, and then watching the guess actually work (or lead the LLM to figure out) is a specific kind of thrill, and it turned out I will chase it for hours…

The best moves I made were judgment, not code. Maybe because I like games, but I was immediately reading the challenge description as a clue instead of a warning, and deciding to stop pushing a dead end and change direction, were my calls, and the AI would not have found the solution without me (and I would not have without the AI).

Legends finished 5th out of 20 teams with 8,797 points and 22 challenges solved. Ruhan solved nineteen of those. I solved three. I assumed this was because he had done this sort of thing many times before, but it turns out it was only his second CTF event.

The top of the board closed tight, two teams tied for first at 9,627 and two more tied for third at 9,127, so fifth was about 830 points off the lead. I don’t think this is last of the Legends team…


Posted

in

, , ,

by