Azure-Oriented Solution Architect Needed
Budget: ₹1,500 – ₹12,500 INR
# Job Title: High-Performance Data Retrieval from XML-Based Patents Database
# Description:
We are seeking an experienced database and data modeling expert to help us optimize the retrieval process from our patents database. Our database is approximately 7TB in size, contains around 160 million records, and is structured across multiple tables with data primarily in XML format.
# Objective:
Our goal is to efficiently fetch 100,000 (1 Lakh) records or more based on their publication numbers within 5-10 seconds.
# Requirements:
1. **Data Structure**:
- The database comprises multiple tables.
- Data is stored in XML format.
- Records are identified by publication numbers.
2. **Performance Goals**:
- Retrieve 100,000+ records within 5-10 seconds.
3. **Technical Challenges**:
- Large database size (7TB).
- XML data processing.
- Multiple tables and potentially complex queries.
# Key Responsibilities:
1. **Database Optimization**:
- Analyze current database schema and indexing strategies.
- Propose and implement indexing improvements tailored for XML data.
- Optimize database queries for speed and efficiency.
2. **Data Modeling**:
- Evaluate the current data model and suggest improvements.
- Implement data modeling techniques that enhance query performance.
3. **Technology Stack**:
- Recommend and implement tools or technologies that can aid in faster XML data processing and retrieval.
- Consider the use of in-memory databases, caching solutions, or distributed databases if applicable.
4. **Testing & Validation**:
- Conduct performance testing to ensure the solution meets the 5-10 second retrieval requirement.
- Provide documentation and training for our team to maintain and further optimize the system.
# Preferred Qualifications:
Proven experience in handling large databases (preferably in the patents or similar domain).
Expertise in XML data handling and optimization.
Strong understanding of indexing, query optimization, and data modeling.
Familiarity with high-performance computing and database solutions (e.g., in-memory databases, distributed systems).
Ability to work independently and deliver results within specified timelines.
# Description:
We are seeking an experienced database and data modeling expert to help us optimize the retrieval process from our patents database. Our database is approximately 7TB in size, contains around 160 million records, and is structured across multiple tables with data primarily in XML format.
# Objective:
Our goal is to efficiently fetch 100,000 (1 Lakh) records or more based on their publication numbers within 5-10 seconds.
# Requirements:
1. **Data Structure**:
- The database comprises multiple tables.
- Data is stored in XML format.
- Records are identified by publication numbers.
2. **Performance Goals**:
- Retrieve 100,000+ records within 5-10 seconds.
3. **Technical Challenges**:
- Large database size (7TB).
- XML data processing.
- Multiple tables and potentially complex queries.
# Key Responsibilities:
1. **Database Optimization**:
- Analyze current database schema and indexing strategies.
- Propose and implement indexing improvements tailored for XML data.
- Optimize database queries for speed and efficiency.
2. **Data Modeling**:
- Evaluate the current data model and suggest improvements.
- Implement data modeling techniques that enhance query performance.
3. **Technology Stack**:
- Recommend and implement tools or technologies that can aid in faster XML data processing and retrieval.
- Consider the use of in-memory databases, caching solutions, or distributed databases if applicable.
4. **Testing & Validation**:
- Conduct performance testing to ensure the solution meets the 5-10 second retrieval requirement.
- Provide documentation and training for our team to maintain and further optimize the system.
# Preferred Qualifications:
Proven experience in handling large databases (preferably in the patents or similar domain).
Expertise in XML data handling and optimization.
Strong understanding of indexing, query optimization, and data modeling.
Familiarity with high-performance computing and database solutions (e.g., in-memory databases, distributed systems).
Ability to work independently and deliver results within specified timelines.