Remote
Full Time
Intermediate or Experienced
Bogotá, Bogota, Colombia
Ciudad de México, Ciudad de México, Mexico
Buenos Aires, Buenos Aires, Argentina
Rio de Janeiro, State of Rio de Janeiro, Brazil
About BOLD Business
About the job
About Bold Business
Bold Business is a U.S.-based, AI-first company building automation systems, AI agents and production software for enterprise and SMB clients. Our internal agents—including Hermes and OpenClaw—perform real operational work across Sales, Recruiting and other business teams.
About the Role
We are looking for an experienced Automation Engineer to build, improve and support production AI agents and integrations.
This is a hands-on role covering agent development, backend automation, data pipelines and system integrations. You will work with our Lead Architect and directly with users to understand their workflows, confirm requirements and deliver reliable solutions.
We want someone who takes ownership while communicating regularly. You should ask questions, educate users, share progress and request feedback—not work alone for long periods without checking in.
What You Will Do
AI Agent Development
- Build and maintain Hermes, OpenClaw and other AI agents using Python.
- Develop multi-agent systems where supervisor agents coordinate worker agents.
- Manage short-term, session-level and long-term agent memory.
- Build tool and function-calling integrations that allow agents to take real actions.
- Test prompts systematically and regression-test important prompt changes.
- Handle agent failures, retries and fallbacks so workflows fail safely.
Integrations and Data
- Connect agents to platforms through APIs, webhooks and event-driven workflows.
- Build webhook receivers that start agent workflows from external events.
- Create idempotent, retry-safe integrations that prevent duplicate actions.
- Maintain an abstraction layer that supports different LLM providers.
- Build data pipelines for ingestion, transformation and validation.
- Manage vector databases used for RAG and long-term agent memory.
- Protect data quality so incorrect data does not reach production agents.
Reliability, Cost and Security
- Make agent workflows traceable and easy to diagnose.
- Log tool calls, LLM responses, agent decisions and failures.
- Control token costs through model routing, caching and reduced duplicate calls.
- Create alerts for failures, high costs and slow performance.
- Protect agents from prompt injection and unsafe input.
- Limit agent permissions so each agent accesses only what it needs.
- Protect personal and sensitive data in prompts, context windows and logs.
User Partnership
Your first users will be our internal Sales and Recruiting teams. You will:
- Learn their workflows, problems and priorities.
- Explain what the technology can and cannot do.
- Confirm requirements and expected results before building.
- Provide regular updates and raise risks or blockers early.
- Demonstrate work and request feedback throughout development.
- Confirm that each solution works in the user’s real workflow.
- Monitor results after launch and recommend improvements.
Required Qualifications
- 5+ years of experience in automation engineering, backend systems or AI development.
- Strong production-level Python skills, including clean code and automated testing.
- Experience building and operating AI agents or LLM-integrated systems in production.
- Understanding of multi-agent coordination and failure handling.
- Production experience with OpenAI, Anthropic, Groq or similar LLM APIs.
- Experience handling API errors, rate limits, monitoring and LLM costs.
- Understanding of short-term, session-level and long-term agent memory.
- Experience building data pipelines for ingestion, transformation and quality control.
- Strong SQL and database-design skills; PostgreSQL experience preferred.
- Experience with REST APIs, webhooks, event-driven systems and idempotency.
- Experience with Pinecone, pgvector, Chroma, Weaviate or another vector database.
- Experience with Git and standard CI/CD practices.
Nice to Have
- LangChain, LangGraph, LlamaIndex or similar frameworks.
- Experience building an LLM provider abstraction layer.
- Make, Zapier or n8n experience.
- AWS or GCP experience.
- Docker and containerization.
- Prompt regression testing or LLM evaluation tools such as RAGAS.
- Familiarity with OpenClaw or similar agent infrastructure.
What Success Requires
- Strong English: You must communicate clearly in writing and on video calls with a U.S.-based team.
- Production experience: You must be able to explain systems you personally helped build and launch.
- Ownership with communication: Move work forward independently while staying aligned with users and technical leadership.
- Accountability: Identify problems early, communicate them clearly and take responsibility for resolving them.
Benefits
Remote Work
Preferred Skills
Python (production-level, clean code, automated testing)
AI Agent & Multi-Agent Architecture (orchestration, Hermes, OpenClaw)
LLM APIs & Frameworks (OpenAI, Anthropic, Groq, LangChain, LangGraph, LlamaIndex)
Agent Memory Management (short-term, session-level, long-term)
Tool & Function Calling Integrations
Prompt Engineering & Testing (regression testing, LLM evaluation tools like RAGAS)
System Failover & Reliability (error handling, retries, fallbacks)
API, Webhook, & Event-Driven Architecture (REST APIs, idempotency, retry-safe systems)
LLM Abstraction Layers
Data Pipeline Engineering (ingestion, transformation, validation, quality control)
Relational Databases & SQL (PostgreSQL, database design)
Vector Databases & RAG (Pinecone, pgvector, Chroma, Weaviate)
Cost Optimization & Monitoring (token management, caching, logging, alerting, cost/performance tracking)
AI Security & Privacy (prompt injection protection, least-privilege access, sensitive data masking)
Version Control & DevOps (Git, CI/CD practices, Docker/containerization, AWS/GCP)
No-Code/Low-Code Automation (Make, Zapier, n8n)
High Ownership & Proactivity
Collaborative & Transparent Communication (regular check-ins, progress sharing, early risk reporting)
Strong English Communication (written and verbal for U.S.-based team collaboration)
User-Centric Mindset (empathy for internal teams, workflow understanding, educating non-technical users)
Accountability & Problem-Solving
Requirement Validation & Feedback-Driven Iteration
Systems-Thinking & Production Mindset