Remote
Full Time
Intermediate or Experienced
Guadalajara, Jalisco, Mexico
Pune, Maharashtra, India
Taguig, Metro Manila, Philippines
About BOLD Business
About the job
Overview
We need a data enthusiast who pairs sharp analytical capabilities with the leverage of modern AI tools. As our organization harmonizes multiple disparate systems, data sources, and analytics platforms into a single operational home, this role will lead and orchestrate the collection, cohesion, and maintenance of our core data models (unified accounts, contacts, candidates, deals, activity, and work records).
You will design and oversee the foundational data layer that allows future AI agents to read seamlessly across all entity data without an integration tax, setting the groundwork for our incoming ERP consolidation.
What You Own
- Canonical Data Model: Define entity structures, relationships, and identity-resolution logic across accounts, contacts, candidates, deals, projects, and work records.
- Validation at Ingest: Establish upfront validation rules that prevent low-quality records from entering systems, shifting focus from post-hoc error reporting to preventative data hygiene.
- Integration Layer Coordination: Collaborate with full-stack engineers to maintain an adaptable integration layer, ensuring vendor replacements require simple configuration rather than a full code rewrite.
- Reporting & Dashboards Foundation: Build and govern the modeled semantic layer beneath our Power BI transition, unifying legacy reports and setting up actionable dashboards for the future ERP.
- Agent Data Access: Architect clean, permissioned, and documented read paths so internal AI agents consume structured data directly rather than scraping front-end source systems.
- Data Quality Ownership: Serve as the direct point of contact and owner for recurring systemic data defects.
- Migration Strategy: Lead data migration paths off retired legacy systems while preserving historical continuity and audit trails.
- Lean AI & Continuous Improvement: Apply Lean principles and AI tools to reduce data failure points, detect early warning indicators, and define KPIs/SLAs to optimize organizational performance.
- Unstructured Data Insights: Monitor, analyze, and process speech, text, and sentiment analysis inputs across business channels.
Requirements
- 4+ years of hands-on experience in data and analytics, featuring heavy data-modeling responsibilities.
- Proven track record of designing a canonical data model across multiple source systems, with the ability to articulate trade-offs in identity resolution and deduplication.
- Deep knowledge of PostgreSQL (schema design, normalization trade-offs, database migrations, query performance tuning).
- SQL proficiency for data exploration, auditing, reconciliation, and cross-database validation.
- Python expertise for developing transformation scripts and data pipelines.
- Demonstrated experience building and supporting integrations with commercial SaaS APIs (handling rate limits, unexpected schema updates, and partial failures).
- Strong understanding of data-layer security, designing models where authorization and field-level restrictions are enforced structurally.
- Proactive adoption of AI tools to accelerate pipeline creation, audit data quality, and enrich datasets.
- Inherent curiosity and rigor for inspecting, validating, and establishing trust in business metrics.
Strong Preferences
- Direct familiarity with core enterprise data models across ERP, HRIS, WFM, or CRM platforms (e.g., NetSuite, SAP, Workday, Salesforce, Genesys, Oracle, or Aspect).
- Experience using tools like dbt (or comparable transformation and automated testing frameworks).
- Hands-on experience working with Tableau, Power BI, or similar analytics stacks, specifically using AI to extract insight from structured datasets.
- Practical experience consolidating legacy commercial systems onto unified or in-house platforms.
- Exposure to recruitment, workforce, or operations data, including handling candidate PII and data retention requirements.
- Prior experience preparing database schemas or data warehouse layers specifically to feed LLMs, AI workflows, or autonomous agent routines.
Benefits
Remote Work
Preferred Skills
Canonical Data Modeling
Python Data Pipelines
PostgreSQL & Advanced SQL
SaaS API Integration
AI & Agent Data Access Architecture