I would agree Walden. A potential approach for this would be generating vector embeddings for each new post (title + first post contents). This post could be compared to other posts by the user within a recent time frame. If they have a high cosine similarity then they are likely a spam post. This could be applied for newer accounts until they reach a certain threshold which shows organic user activity.
The cost for this would be very minimal as text embeddings typically bill at 0.02 per 1M input tokens. A simple formula for this is:
Cost = ( new accounts per day ) • (tokens per post) • (.02/1M)
(10,000) • (420 ) • (.02/1M)
= $0.084 (obviously this assumes 1 post per bot, but I'm overestimating on new user signups probably)
False positives could be reduced by
- Requiring more than X posts to be in violation.
- Checking semantic similarity with typical rock climbing posts. Ie. Is it about rock climbing at all?
- Using TF/IDF as an embedding component for the site. We can better represent/factor in each sub-fourm's discussion nuances.
- Comparing the embeddings of a user to known past spammers.
Idk, I don't work in this space and I'm probably over simplifying/not factoring in a lot of nuances, but vector embeddings are sweet and they seem well suited to typical spam posters I've seen in the past few months.