Founding ML Engineer — Live Music Intelligence
Gotobeat – The operating system for live music
Location: Remote-first (UTC to UTC+3 preferred) | Type: Full-time | Reports to: Head of Tech
Salary: Competitive + equity
How to apply: send your CV to alfredo@gotobeat.com
The problem
Can this artist fill this room, in this city, on this Thursday? Every promoter bets their own money on that question, and most of them answer it with gut feel, a Spotify listener count and a phone call to a friend. A good answer puts the right act in a 300-cap room in Leeds two summers before it headlines a festival. A bad answer leaves a half-empty venue and a lost deposit.
We built a system that answers it better than gut feel. It is called Antonello. We are hiring the person who will own it.
About Gotobeat
Gotobeat builds software for live music: for the artists who play, the venues that host and the promoters who book. We help emerging artists earn more, venues fill more seats and promoters spend less time on admin. The platform covers planning, promotion, ticketing and post-show settlement. It runs on an AWS serverless stack, Python data pipelines and LLM agents.
Why this role matters
Most music data companies see one slice: streams, or social, or ticket sales. Gotobeat is a ticketing platform and a promoter tool, so we hold the dataset that ties them together: what an artist streamed, what they posted, where they played and how many people paid at the door. That join is the most valuable thing we own. Today one founder maintains it.
You are the first hire whose whole job is that data. You take over the models, the pipelines and the storage, and you own them end to end: from the notebook where a feature starts to the number a promoter reads before they sign a deal. When your forecast moves, shows get booked or not. The feedback is an empty room or a full one.
What you take over on day one
- Antonello, the capacity engine. A per-city venue-capacity forecast: an XGBoost ensemble over an event-history prior, calibrated ranges, confidence gates that hold back a weak prediction, and an LLM evidence-repair loop that finds the shows our scrapers missed.
- The discovery loop. A daily pipeline that pulls new artists from Songstats, Spotify playlists and related-artist graphs, runs them through fit gates, genre families and liveness checks, and gives a promoter a shortlist that matches their taste and their rooms.
- The data estate. DynamoDB as the system of record, a change-data-capture mirror into a normalised Postgres schema on Supabase, S3 snapshots, and the API clients and data-partner integrations for streaming, social and touring data.
- The analytics. Sell-through and price suggestions, the hype and momentum scores, ad attribution across Meta and Google, and a rule engine that checks the model when the evidence is thin.
What you'll do
- 35% Build the best artist model in the music industry. Antonello forecasts what an act can draw in a city. Your job is to make that the number the industry trusts. Retrain it, recalibrate it and score it against what happened at the door. Decide when it should stay silent. Ship each improvement to production yourself.
- 25% Own the data pipeline and the ingestion. Keep every source complete, fresh and cheap, and make the join between streaming, social, touring and ticket sales something a colleague can query in one line.
- 15% Own the architecture of the data estate: the single-table design, the change-data-capture mirror into Postgres, and the schema every model and every query reads. Design the contracts so a new source or a new model plugs in without a migration weekend.
- 15% Build the analytics a promoter bets on: expected sell-through, the right ticket price, which ad spend sold tickets, and the confidence behind each number.
- 10% Expose the models as tools our LLM agents call, with the evaluation and the guardrails that let an agent draft a deal from a forecast.
The stack
- Python 3.13 in Lambda containers for the models and the scrapers: XGBoost, numpy, Playwright.
- TypeScript on Node for the orchestration, the API and the promoter product: AWS Lambda, Step Functions, EventBridge, and SST v3 as infrastructure as code.
- DynamoDB with ElectroDB, DynamoDB Streams into Supabase Postgres, S3 for snapshots and artefacts.
- Claude, Gemini and Perplexity for the agents, called through MCP tool surfaces.
- No Kubernetes, no Spark cluster, no GPU farm. The data is small: tens of thousands of artists, hundreds of thousands of shows and the ticket sales behind them.
- No pager. The crons run overnight and the alarms land in Slack.
You'll be successful here if you have
- Substantial experience of shipping machine learning that people relied on, including at least one model that was wrong in production and what you learned from it.
- Depth in tabular ML and forecasting: gradient boosting, calibration, evaluation on small and noisy data, and a preference for the simpler model that you can explain to a promoter.
- Data engineering: you have designed a schema, owned a pipeline end to end and made an expensive source cheap.
- Fluent Python, and enough TypeScript to ship the glue yourself.
- Practical LLM engineering: tool use, retrieval and evaluation, and a clear view of when a model must ask before it acts.
- Experience of building on your own: from a blank repo to something people use, with nobody else writing the tickets.
- A love of live music. You have stood in a half-empty room and known why.
Nice to have
- Experience of music or entertainment data, for example at a streaming service, a data provider, a ticketing platform, a label or a booking agency.
- Experience with streaming, social or touring APIs, and with an ingestion that survives an upstream change.
- Experience with AWS serverless data patterns and change-data-capture pipelines.
- Experience with the MCP tool protocol and agent evaluation.
How we work
- Remote-first, with quarterly in-person hack weeks in London.
- A small, senior team with no management layers and a lot of autonomy.
- One-week sprints, async communication and fast decisions.
- The founder who built the current models works with you for the first months, then hands over.
- We value clean code, continuous learning and doing right by artists.
Benefits
- Competitive salary, set by experience, plus founding-team equity.
- Fully remote, with the hours that suit you.
- We engage you as an employee or as a contractor, according to the country where you live.
- Unlimited paid holiday, plus local public holidays.
- Free gig tickets and backstage passes to all Gotobeat shows.
- A budget for the hardware, the data and the conferences you need.
Hiring process
- Intro chat (30 min): the role, the music and whether we make sense to each other.
- Technical deep dive (90 min): you walk us through a model you shipped, then we pair on a prediction from our data that went wrong. No LeetCode, no take-home.
- Founders' chat (30 min): vision, values, equity and any remaining questions.
Total time from first call to offer: about two weeks.
Ready to join Gotobeat?
Send your CV to alfredo@gotobeat.com with the subject line "Founding ML Engineer – [Your Name]". Tell us about a model you shipped and what it got wrong, and name one artist you think is about to break. We reply within five business days.
Gotobeat is an equal-opportunity employer. We welcome applications from everyone. Tell us if you need any adjustment to the process.
We use your application only to assess you for this role. We keep it for six months after the role closes, then we delete it. See the privacy policy for your rights.
See the other open positions on the Gotobeat hiring page.