Tuesday, March 24, 2009 

Data Mining Survey - Last Call

Rexer Analytics has just issued a last call for its annual data mining survey. This is a pretty nice survey that provides a great deal of valuable information about how data mining is used and who is doing it.

To participate, please click on the link below and enter the access code in the space provided. The survey should take approximately 20 minutes to complete.  At the end of the survey you can request to have the results sent to you as well as get a copy of last year's survey.

Survey Link: www.RexerAnalytics.com/Data-Miner-Survey.html
Access Code: RS2008

Labels:

Tuesday, October 28, 2008 

Oracle BIWA Summit 2008


The Oracle BIWA Summit 2008 is approaching  (December 2-3) . It will be held at Oracle World HQ, Redwood Shores, California. This is the second event of its kind. Last year's event was a great success and lots of fun (see details here ). This year's keynotes include Jeanne Harris (co-author of "Competing on Analytics") and Usama Fayyad (legendary data miner).  Here are some information and links about the event:

Sign up now to attend the Oracle BIWA Summit 2008 Dec. 2-3.  Attend this unique two-day IOUG Business Intelligence, Warehousing and Analytics (BIWA) SIG (www.oraclebiwa.org) event to gain the knowledge and information critical for success in your work.  Attend 65 technical talks and 10 hands-on sessions, hear keynotes from Jeanne Harris, the co-author of the best-seller Competing on Analytics, and other industry leaders, learn the latest trends in data warehousing, business intelligence and analytics best practices, learn how to overcome common challenges and network with your peers.

Learn how to improve data warehouse query performance by a factor of 10x with Oracle Exadata and hear firsthand from Oracle Senior Executives and other experts about the revolutionary new HP Oracle Exadata Storage Server and HP Oracle Database Machine and how they fit into Oracle’s data warehousing strategy.

Keynotes
Technical Talks and Hands-on Sessions  also see BIWA Summit agenda
Travel and Logistics  also see BIWA Summit invitation
Sponsors
REGISTER for BIWA now!  - Last year's BIWA Summit 2007 was sold out, so sign up now to reserve your space.

BIWA Summit 2008
December 2 - 3, 2008





Labels:

Friday, August 15, 2008 

Oracle at KDD 2008 and KDD 2007 Pictures

It is that time of the year again. In about a week I am going to be attending the KDD (Knowledge Discovery in Databases) 2008 conference (conference website) along with some other Oracle colleagues. KDD is one of the primary conferences on data mining. This year it will take place in Las Vegas, Nevada, from August 24 to 27.

Oracle is a Gold sponsor for the event and will have again a large presence at the conference. Among other things, Oracle is the sponsor of the first poster reception and demonstration session on Monday.

This year I will not try to post daily conference highlights. After managing to do a poorer job last year than in the year before I gave up on the idea. Instead I will settle for a single post after the conference.

Here are some pictures from the Oracle booth at last year's KDD:


If you plan to attend KDD this year, stop by the Oracle booth to say hello. Last year I spent a great deal time at the Oracle booth. There were lots of interesting questions and problems. Here are some of the topics we talked about:

  • How to combine data mining with a rules engine?
  • How to schedule a background process for anomaly detection in tables?
  • How to use database triggers to score new rows or modification of rows?
Because Oracle has a large array of features in the RDBMS, it is possible to come up with very interesting solutions to these and other questions by combining these features.

Labels: ,

Wednesday, May 02, 2007 

Webcast Announcement: Oracle's In-Database Statistics

Today (Wednesday), May 2, 2007 at 12:00 PM EST, the Oracle Business Intelligence, Warehouse and Analytics (BIWA) Special Interest Group (SIG) will host another interesting free webcast:

Oracle's In-Database Statistics
Speaker: Charlie Berger

Session Abstract
Oracle Database 10g embeds a range of SQL-based basic statistical functions including: summary statistics, hypothesis testing, correlation coefficients, distribution fitting, ANOVA, and cross-tabulation statistics. This collection of useful in-database statistical functions offer an attractive alternative to SAS and other external statistical packages. Additionally, in-database statistical functions improve "information latency" and enable analytical processes to be performed without the traditional complicated "data extraction, analysis and import of results" process. This BIWA eSeminar will introduce Oracle's statistical functions using examples.

  • Perform summary statistics (mean, median, quartiles, extreme values, etc.)
  • Perform various hypothesis tests (t-test, F-test, Mann Whitney test, ANOVA, etc.)
  • Perform correlations analysis, cross tabs, etc.
  • Perform distribution tests
Speaker
Charlie Berger is the Senior Director of Product Management for Data Mining Technologies at Oracle. He has over 20 years of experience in data mining and statistical software industry including product management positions at Thinking Machines Corporation and Bolt Beranek and Newman (BBN) Corporation. He holds Masters in Manufacturing Engineering and MBA degrees, both from Boston University.

Web Conference link
Video link: Click here.

Audio link (USA): 888-967-2253
Meeting Id: 534705#
Passcode: 334451#
Audio will also be streamed on web for overseas participants.

If you have problems connecting using the above information check for updated web conference information at the BIWA web page.

BIWA SIG
BIWA SIG is a forum where customers can exchange the latest information on Oracle BI/DW technologies. The BIWA SIG regularly invites speakers to give talks of potential interest to its members. Visit BIWA for more information.

Labels:

Tuesday, April 24, 2007 

Webcast Announcement: A Simple Fraud Detection Application using Oracle Data Mining, SQL Developer and Oracle BI EE

Tomorrow, April 25, 2007On April 25, 2007 at 11:45 AM EDT, the Oracle Business Intelligence, Warehouse and Analytics (BIWA) Special Interest Group (SIG) will host the following free webcast:

A Simple Fraud Detection Application using ODM, BIEE, and SQL Developer
Speaker: Bob Haberstroh

Session Abstract
Classification is an often-used methodology in data mining that creates a predictive model distinguishing between (or among) two or more classes of individuals - for example: which of my customers is likely to churn (or not). However, in some cases there is insufficient evidence to profile one of the classes, such as in cases of criminal or fraudulent activity. Oracle Data Mining includes an algorithm for Anomaly Detection that uses the so-called “One Class Classifier” technique to identify rare occurrences of likely abnormal or suspicious activity.

This presentation gives an overview of this functionality and includes a live demonstration of Anomaly Detection using Oracle Data Mining. The methodology of automating the process in an application will also be indicated.

Speaker
Bob Haberstroh is Principal Product Manager for Data Mining Technologies at Oracle. He has many years of experience in statistics, and data mining including more than 12 years in product management positions at Thinking Machines Corporation and Oracle.

Web Conference link
Video link: Click here.

Audio link (USA): 888-967-2253
Meeting Id: 534705#
Passcode: 334451#
Audio will also be streamed on web for overseas users.

If you have problems connecting using the above information check for updated web conference information at the BIWA web page.

BIWA SIG
BIWA SIG is a forum where customers can exchange the latest information on Oracle BI/DW technologies. The BIWA SIG regularly invites speakers to give talks of potential interest to its members. Visit BIWA for more information.

Labels:

Friday, December 15, 2006 

Announcement: Oracle Data Mining Consultants Partnership Program

We're starting a program to work with qualified data mining consultants.

You and your colleagues are invited to participate in a 2 day hands-on session designed for data mining consultants here in the Oracle Burlington MA office February 7 & 8, 2007. It is also possible to attend remotely via webminar. Space is limited, so please RSVP asap.

The Oracle Data Mining Consultants Partnership Program has been established to develop a support network of skilled data mining consultants. As part of this program, we're looking to promote knowledgeable, skilled data mining experts who can leverage Oracle's in-database functionality. Consultants that demonstrate proficiency in Oracle Data Mining and working in Oracle-centric environments will be promoted within Oracle Sales via eSeminars, web directories, user organizations, and face to face meetings.

Assuming that you already know how to "mine" data, the objective of the 2 day training would be to get you comfortable using Oracle Data Mining. We also want to stress deployment, SQL and Java APIs, Predictive Analytics (one-click data mining PL/SQL packages), text mining & Oracle's in-database statistical functions.

Agenda

Wed, Feb 7

8:00-9:30 Arrive early for Installation Assistance

For those who don't already have Oracle Database 10gR2 with the Oracle Data Mining Option, Oracle Data Miner, SQL Developer & JDeveloper, and for added bonus, Oracle BI EE already installed on their machines

9:30-12:30 Hands-on Workshop

Step by step guide using Oracle Data Miner and the Oracle Data Miner Tutorial for data preparation, data transformations & ODM algorithms using various demo data

12:30-1:30 Lunch & Networking

With Oracle Data Mining Developers & Product Management & Sales Management

1:30-5:30 Hands-on Workshop

Step by step guide using Oracle Data Miner for data preparation, data transformations & ODM algorithms using various demo data. (Start on Day 2 material if finish Day 1 early)

6:30 Group dinner

Thurs, Feb 8

9-12:30 Applications Development, Deployment & ODM's SQL & Java APIs

Walking through sample programs and applications use cases

12:30-1:30 Working Lunch

Exposure to ancillary Oracle DWBI&A products (OWB, Oracle BI EE, etc.) 20 min presentations

1:30 - 3:30 Applications Development, Deployment & ODM's SQL & Java APIs

Walking through sample programs and applications use cases

3:30 - 5:30 Oracle Data Mining 11g Preview & Feedback

It requires signed non-disclosure form

5:30 - 6:00 Feedback & Wrap up

You will be expected to bring your own PC but will be able to download, install and use Oracle software (for evaluation purposes). Additionally, you may want to consider joining Oracle's Partner Network which extends many benefits beyond the Oracle Data Mining Consultants Partnership Program.

Please RSVP to Evelyn Farrell ASAP as space is limited.

Hopefully see you in February!

Edit: I forgot to mention that it is also possible to attend remotely. I've added information about this.

Labels:

Tuesday, October 31, 2006 

Free Webinar: Competing on Analytics

I blogged some time ago (link) about an article on The Harvard Business Review by Babson College's Tom H. Davenport on how analytics are becoming a key competitive factor for companies. I have just learned that Prof. Davenport is giving a free webinar today. The theme is "Competing on Analytics." What participants will learn:

  • What data-driven marketing is (and isn't)
  • How marketing visionaries like Capital One, P&G, Amazon, and the New England Patriots are using analytics for competitive advantage
  • What specific tactics these early adopters believe are essential to their success (and what they'd do differently next time)
  • How you can personally succeed as a marketer during these tumultuous times
Check it out. The webinar is Today (October 31, 2006) at 10:00 AM PST/ 12:00 PM, CET/ 1:00 PM EST. Go to this link to register. Again, it is free.

Labels:

Monday, August 21, 2006 

KDD 2006 - Day One

Philadelphia city hall seen from the hotelKDD concentrates most of the tutorials and workshops on the first day. In previous years I usually jumped around from room to room trying to catch interesting talks. This year I decided to follow a different strategy. I picked a full day workshop and stuck with it for the day. I chose the Data Mining for Business Applications Workshop organized by Rayid Ghani (Accenture Technology Labs) and Carlos Soares (University of Porto). It seems that I picked well. The meeting, in one of the larger rooms, was quite full and the audience was very participative. The panel discussions were very lively. I also got to hook up with the rest of the Oracle team and got an early look on our presence on the exhibit hall. Pretty cool stuff, more on that on the next post.

There were many interesting talks and discussions. I will summarize the high points for me. The first talk I saw, A Boosting Approach to Automated Trading (Creamer & Freund), described an automated approach for trading using machine learning. The results were very good. One of the questions from the audience captured something that always comes to mind when I see a talk like that: If this system is so good why is someone talking about it instead of making money with it? The answer came at many levels. First there is the desire to make public one's research. Second the presented system was a simplified system. A real system would have higher scalability requirements. But the key issue was a matter of trust. Would you risk your money on an automated machine learning approach? Trust was a theme that recurred during the whole workshop.

The next talk, A Decision Management Approach to Basel II Compliant Credit Risk Management (van de Putten, et al.), stressed the need to combine machine learning/data mining approaches with rule-based ones that codifies knowledge from the user. The combination of data mining and rules are essential for success. A key role played by data mining in this context is the estimation of the probability of default on a loan. This can have a very big impact on how much money banks are required to set aside under the Basel II rules. This seems to be an interesting problem for trying a combination of Oracle technologies, namely: Oracle Data Mining and Oracle's rules engine technologies. The speaker proposed many areas for future research. When I asked him about where to find real data for meaningful research he acknowledged that there are no publicly available data sets. I see this as one of the key factors preventing the development of applications using data mining. Without meaningful and realistic data sets publicly available, research progresses very slowly. Only a small number of researchers that have access to proprietary data can make contributions. The data mining research community at large is left untapped as a resource to find solutions to pressing problems. Furthermore, if realistic data were available, it would be easier to develop prototypes that showcase the value of the technology to end-users in a compelling way. This could have a very positive impact in increasing the adoption of analytics in applications. The lack of available data was another recurrent topic in the workshop.

Another interesting talk was Discovering Telecom Fraud Situations Through Mining Anomalous Behavior Patterns (Alves, et al.). The authors used clustering as the modeling technique. I wonder if better results cannot be obtained using a technique like one-class support vector machines included in Oracle Data Mining. Again the data is not publicly available.

The talk Interactivity Closes the Gap: Lessons Learned in an Automotive Industry Application (Blumenstock, et al.) discussed a human in the loop approach. Users interact with a data mining tool to gradually build a solution that they trust. During the discussion the issue of accuracy vs. transparency came up. Practitioners feel that in the earlier stages of a project transparency considerations dominate (the trust issue again). In later stages, as the user starts to trust the technology, accuracy becomes more important even if transparency is lost.

The panel discussions were very lively. The first panel addressed Bridging the Gap Between Data Mining Research and Practical Business Applications (Ronny Kohavi, Karl Rexer, and Galit Shmueli). Kohavi contrasted his experiences at Amazon and Microsoft. Two companies with very different styles regarding research. Research at Amazon is very focused and secretive. Research at Microsoft has a broader scope and is very open. Amazon approach makes it easier to turn research into products. Microsoft approach is better at recruiting talented researchers. Shmueli proposed that MBA students are good material for data mining recruitment. She has also observed an increase in the number of MBAs taking courses in analytics. Rexer commented that he works with many large companies that do not have a dedicated analytical group. He proposed that we need to focus training on how to use data mining in business instead of data mining per se. I agree with that. There is too much focus on techniques and very little discussion on business uses.

The second panel discussion, Deploying Data Mining Solutions: Stories, Challenges, and Open Issues (Tyler Kohn, Ramin Mikaili, Richard Boire, and Françoise Fogelman) covered a number of interesting use cases. Mikalli presented a framework used by Accenture to solve problems with good success on some challenging problems. Boire showed how a series of real-life business concerns were addressed with data mining. He also highlighted how easy is to obtain incorrect results when one is not careful applying data mining techniques. Fogelman presented the most controversial talk. She articulates the need for automating data mining in order to empower a larger number of users and eliminate the analytical bottleneck. I liked here talk quite a bit as I am a strong believer that automation of data mining is not only a necessity but unavoidable. However, automating data mining is a hard sell to an audience of data mining experts. The topic generated a very animated discussion. The usual reaction to automation is to list a number of situations where it would be very hard to do it. But that is to miss the point. There will always be problems that need experts. Automation allows users to solve the simpler problems by themselves. It frees the experts to solve the hard problems. As a result the overall utilization of analytics in a companies increases. This also makes the value of the technology more tangible to end users. As users better understand the value of data mining, the number of jobs available to experts should also increase (remember that there are always tough problems to solve).

Labels: ,

Friday, February 17, 2006 

Oracle Life Sciences Meeting

The OLSUG Conference Announcement & Agenda for the Oracle Life Sciences User Group Meeting in Boston on April 3 are now posted on the Oracle Life Sciences User Group and the OTN Life Sciences web sites.

The OLSUG Conference Program includes three tracks of industry leaders, technical experts, Oracle experts, and a Hands-on Technical Workshop with 30 PCs loaded with Oracle 10g Release 2 and a number of technical demos including: Data Mining, Statistical functions, RDF and the Semantic Web, InterMedia & Images, HTML DB (now called Applications Express), Text Mining of Medline, BLAST, JDeveloper, and other useful demos for life sciences & healthcare applications.

This is a great opportunity for customers, prospective customers and Oracle personnel to share, exchange, and experience "best practices" in the Life Sciences & Healthcare industry.

Labels:

About me

  • Marcos M. Campos: Development Manager for Oracle Data Mining Technologies. Previously Senior Scientist with Thinking Machines. Over the years I have been working on transforming databases into easy to use analytical servers.
  • My profile

Disclaimer

  • Opinions expressed are entirely my own and do not reflect the position of Oracle or any other corporation. The views and opinions expressed by visitors to this blog are theirs and do not necessarily reflect mine.
  • This work is licensed under a Creative Commons license.
  • Creative Commons License

Email-Digest



Feeds

Search


Posts

All Posts

Category Cloud

Links

Locations of visitors to this page
Powered by Blogger
Get Firefox