Solving practical problems with AI – Part 1 – Jpeg files renaming using EXIF data

I don’t write about AI much, because everyone else does. I use it though, more and more, despite being a dedicated late adopter of new technologies, by choice.

I thought it would be interesting to post examples of various practical (often cyber-related) problems, that often require some scripting / programming work done to be solved. Problems that in the past would take a few good hours of work to research, iteratively develop an early prototype script code, test it, troubleshoot it and finalize it so it can be successfully used to solve that given problem. And then I would post it on this blog. And yes, in the past I posted many manually written scripts here, and they were always a result of my many human cycles that I spent on creating them…

It turns out, today many of these practical problems can be solved by AI within less than 5 minutes…

I recently got a few hundred JPEG unsorted image files that I wanted to sort by the time of their creation. My solution idea was simple: add a prefix to all these files using a ISO 8601-like timestamp that I would extract from the EXIF data of each file… In the past a ‘simple’ idea like that would lead me to look at EXIF specification, jpeg file format, existing jpeg/exif handling python/perl modules, and of course, potentially many existing solutions to this known problem, including many code snippets all over the place that I could quickly adapt to my needs.

In 2026 all of this potential work was replaced by a simple ChatGPT prompt:

write a python script that enumerates all image files in the directory, recognizes jpeg files, extracts exif info and uses exif timestamp to rename the file to a format yyyy-MM-dd_hh-mm-ss_ followed by the old file name

The script generated by this prompt not only worked ‘out of the box’ – that is, worked like a charm and helped me to rename the image files the way I wanted. It actually included some code that I didn’t expect. For instance, it handled 3 different exif timestamps “DateTimeOriginal”, “DateTimeDigitized”, “DateTime” that are associated with ‘image creation’ event. Secondly, it handled cases where the input file has been already a subject to a previous run of the script (avoiding adding multiple timestamp-based prefixes). It also made sure it doesn’t overwrite files if the ‘newly generated filename’ was identical with an existing file. Pretty clever. Doing far more than I would if it was my usual quick&dirty scripting exercise. And lo, and behold, it included comments and handled help/dry-run command line arguments too.

The AI technology can be seen as an amazing catalyst, prototyping and rapid development booster and performance multiplier. This can be true.

BUT

You need to know what you want. And for that, you need to know the foundations.

AND

You need to be able to check what you get (what AI produces). I actually did a full code review of the generated script before I executed it the first time. This is to make sure there are no surprises.

The boring state of stalled timelines…

The subject of timelines & supertimelines used to dominate many discussions in the digital forensic world… 10-15 years ago. At that time the concept of temporal proximity was super hot and for a good reason – once you put events in a chronological order you can quickly draw conclusions and easily explain what happened.

Today this topic is kinda dead because many modern forensic and EDR/XDR tools do a great job democratizing and commoditizing the concept of timelines, and these simply became a part of what we can call ‘the business as usual’.

BUT

I believe there is still more to it, so much more that I decided to write a longer post about it.

In 2012 I described a Windows file system-based PE files clustering technique that relies on PE file compilation timestamps to detect suspicious executables. While the newer Windows versions (10+) use the very same PE file field for a different purpose (reproducible builds) there is still some merit to this technique for non OS PE files found on the system.

In 2015 I introduced a new analysis technique that I called filighting. The idea relies on targeted file content analysis focused on installed software packages. The assumption being that all the files belonging to the (atomically) installed software somehow reference each other, at least once. My hypothesis was that by mapping these internal cross-references one can quickly find outliers (possibly malicious files residing in the software’s directory; potential supply chain-attacks). A few follow-up posts demonstrated a visual representation of inter-file connections found in many popular software packages, rendered with a help of d3 library.

10-15 years ago forensic analysis relied on running many separate tools to extract the content/data of many forensic artifacts, individually. Today’s forensic tools often take care of these in an uniform way and simply parse and extract data + put these extracted information pieces on a timeline… This is all cool, but it doesn’t take into account the complexity of today’s ecosystems.I want forensic tools to start clustering these atomic data points a bit more.

The presence of EDR’s telemetry, the .bash_history files and their backups, the auditd and sysmon logs are now a norm. We can also easily take snapshots of file system’s metadata for Windows, Linux and macOS. What these snapshots often tell us in 2026 is that unlike in 2000s, many systems today are often set up not from the scratch but by leveraging massive ‘copy events’ where old data from an old file system X is being blindly copied to a new file system Y. Such activity introduces a lot of challenges to digital forensic professionals who see file system-based timelines that are often massively distorted.

Additionally, many (primarily) Linux systems hosting web sites end up (over time) with many copies of the same website content spread across many similarly looking directories. A proper (and automatic) timeline analysis should detect these clusters of similarly-looking files spread across many ‘backup’ directories. Such analysis may help to immediately highlight files added / changed in different iterations of the same directory (e.g. web shells or files uploaded via file upload vulnerability).

There is also the issue of systems/endpoints heavily utilized by legitimate, authorized internal security teams. Analysis of such systems pose a huge challenge as many activities observed on these systems usually correspond with legitimate activities of internal soc, pentesting, red team operators.

I guess the point of this post is that our timelines need more juice.

The EDR/XDR telemetry is very metadata-centric, but DFIR-access level gives us all the content we need. So, when a forensic software is analysing the file system I want it to find more than just known-knowns. This is the easy bit. I want it to explore use cases where someone copied many files, someone created a copy of another directory, where potential TA uploaded web-shells, even if they cannot be executed, where company employees hone their tradecraft, run offensive tools, do stupid but explainable stuff, where they bypass security controls, harmlessly download and analyze malware, use bad tools for a good reason, and so on and so forth.

I want timelines to be presented as clusters of activity, with attribution and with perceived intent.