Data Warehousing and Data integration

Question about extracting data from PDFs using Gemini

More
5 months 2 weeks ago - 5 months 2 weeks ago #25947 by Neil
Hi Team,

I watched your demo on extracting data from PDFs using Gemini - very impressive.

My main concern is sending data to Gemini (or any AI service) when that data could be considered corporate intellectual property, especially given the way AI models may learn from submitted content.

What are your thoughts on this from a data security and IP perspective?

I realize I could extract the required data from PDFs using a Python library instead, but in the past I’ve found calling Python from within AETL to be relatively heavy.I’d appreciate your feedback.

Thanks in advance.
Last edit: 5 months 2 weeks ago by Neil.
The following user(s) said Thank You: Peter.Jonson

Please Log in or Create an account to join the conversation.

More
5 months 2 weeks ago - 5 months 2 weeks ago #25950 by Peter.Jonson
Hi Neil

My only concern is pushing data to Gemini or any AI that may be considered corporate Intellectual Property especially given the 'learning' that the AI engines go through.

>>> What are your thoughts? 

Well, you have to discuss it with your customers there is no escape from it.

>>> I know I can probably extract the data I need from these PDFs using a Python package but in the past calling Python from with AETL was quite heavy 

We ended up use python for the processing financial statements  (I am the one who wrote the script) Financial statement is a list of invoices EG: I upload my invoices you upload your invoices  and we use web portal for reconciliation.

Although is sounds very simple but every customer is using a slightly different format and naming convention.
Financial statement can be text file, excel file or pdf plus it can span over multiple pages or sheets.

Gemini has got much better but still it hallucinates EG: you submit same file 10 times it works fine than suddenly it gives incorrect result.Right now it is semi automated process. 

AI does not work correctly all the the time  so end user still has to check results manually before loading them.    

We did put a lot of effort into it. 

Some pdf have no text inside they are just images inside the file. 

What we ended up doing is converting pdfs and excel sheets into image files and submitting them to Gemini for processing.

Large files may take 2-3 minutes to process. AETL does all prep/post processing, python converts files to images and submit them to Gemini.

In my opinion, it is very important not to promise too much to the customers and tell them about hallucinations 😎.

I hope it is helpful to you

Peter Jonson
ETL Developer
Last edit: 5 months 2 weeks ago by Peter.Jonson.
The following user(s) said Thank You: Neil

Please Log in or Create an account to join the conversation.

More
5 months 3 days ago #26000 by Neil
Hey Peter.

Thanks for the Insight. Really appreciate  it and I will be sure to keep things transparent for the client.
Your solution still looks good. But yes, Humans need to be there to police it.

Please Log in or Create an account to join the conversation.

Cookies user preferences
We use cookies to ensure you to get the best experience on our website. If you decline the use of cookies, this website may not function as expected.
Accept all
Decline all
Read more
Marketing
Set of techniques which have for object the commercial strategy and in particular the market study.
Google
Accept
Decline
Analytics
Tools used to analyze the data to measure the effectiveness of a website and to understand how it works.
Google Analytics
Accept
Decline
Google Analytics
Accept
Decline
Functional
Tools used to give you more features when navigating on the website, this can include social sharing.
Advertisement
If you accept, the ads on the page will be adapted to your preferences.
Google Ad
Accept
Decline
Save